Kaiser Sun scite author profile

Kaiser Sun

4Publications

8Citation Statements Received

205Citation Statements Given

How they've been cited

How they cite others

166

205

Affiliations

University of Washington

Publications

Order By: Most citations

Effective Attention Sheds Light On Interpretability

Sun¹,

Marasovi²

2021

View full text Add to dashboard Cite

An attention matrix of a transformer selfattention sublayer can provably be decomposed into two components and only one of them (effective attention) contributes to the model output. This leads us to ask whether visualizing effective attention gives different conclusions than interpretation of standard attention. Using a subset of the GLUE tasks and BERT, we carry out an analysis to compare the two attention matrices, and show that their interpretations differ. Effective attention is less associated with the features related to the language modeling pretraining such as the separator token, and it has more potential to illustrate linguistic features captured by the model for solving the end-task. Given the found differences, we recommend using effective attention for studying a transformer's behavior since it is more pertinent to the model output by design.

show abstract

State-of-the-art generalisation research in NLP: A taxonomy and review

Hupkes¹,

Giulianelli²,

Dankers³

et al. 2022

Preprint

View full text Add to dashboard Cite

The ability to generalise well is one of the primary desiderata of natural language processing (NLP). Yet, what 'good generalisation' entails and how it should be evaluated is not well understood, nor are there any common standards to evaluate it. In this paper, we aim to lay the groundwork to improve both of these issues. We present a taxonomy for characterising and understanding generalisation research in NLP, we use that taxonomy to present a comprehensive map of published generalisation studies, and we make recommendations for which areas might deserve attention in the future. Our taxonomy is based on an extensive literature review of generalisation research, and contains five axes along which studies can differ: their main motivation, the type of generalisation they aim to solve, the type of data shift they consider, the source by which this data shift is obtained, and the locus of the shift within the modelling pipeline. We use our taxonomy to classify over 400 previous papers that test generalisation, for a total of more than 600 individual experiments. Considering the results of this review, we present an in-depth analysis of the current state of generalisation research in NLP, and make recommendations for the future. Along with this paper, we release a webpage where the results of our review can be dynamically explored, and which we intend to update as new NLP generalisation studies are published. With this work, we aim to make steps towards making state-of-the-art generalisation testing the new status quo in NLP.

show abstract

Tokenization Consistency Matters for Generative Models on Extractive NLP Tasks

Sun¹,

Pan²,

Zhang³

et al. 2022

Preprint

View full text Add to dashboard Cite

Effective Attention Sheds Light On Interpretability

Sun

Marasović

2021

Preprint

View full text Add to dashboard Cite

scite is a Brooklyn-based organization that helps researchers better discover and understand research articles through Smart Citations–citations that display the context of the citation and describe whether the article provides supporting or contrasting evidence. scite is used by students and researchers from around the world and is funded in part by the National Science Foundation and the National Institute on Drug Abuse of the National Institutes of Health.

Contact Info

hi@scite.ai

10624 S. Eastern Ave., Ste. A-614

Henderson, NV 89052, USA

Blog Terms and Conditions API Terms Privacy Policy Contact Cookie Preferences Do Not Sell or Share My Personal Information

Made with 💙 for researchers

Part of the Research Solutions Family.

Kaiser Sun

Effective Attention Sheds Light On Interpretability

State-of-the-art generalisation research in NLP: A taxonomy and review

Tokenization Consistency Matters for Generative Models on Extractive NLP Tasks

Effective Attention Sheds Light On Interpretability

Contact Info

Product

Resources

About