Contextual Embeddings: When Are They Worth It?

Arora, Simran; May, Avner; Zhang, Jian; Ré, Christopher

doi:10.48550/arxiv.2005.09117

Cited by 8 publications

(9 citation statements)

References 7 publications

(10 reference statements)

Supporting

Mentioning

Contrasting

Order By: Relevance

“…Therefore, the same token in different sentences will have a different embedding. This attention to context correlates with high reliability in several natural language processing tasks, such as named entity recognition, concept extraction, and sentiment analysis, and is relatively better than non-context embedding (Taillé et al, 2020; Arora et al, 2020). In addition, the form of the token, which is part of the sentence, makes BERT more adaptive to typographical errors and variations of word writing.…”

Section: Methodsmentioning

confidence: 99%

CASBERT: BERT-Based Retrieval for Compositely Annotated Biosimulation Model Entities

Munarko

Rampadarath

Nickerson

2022

Preprint

View full text Add to dashboard Cite

Maximising FAIRness of biosimulation models requires a comprehensive description of model entities such as reactions, variables, and components. The COmputational Modeling in BIology NEtwork (COMBINE) community encourages the use of RDF with composite annotations that semantically involve ontologies to ensure completeness and accuracy. These annotations facilitate scientists to find models or detailed information to inform further reuse, such as model composition, reproduction, and curation. SPARQL has been recommended as a key standard to access semantic annotation with RDF, which helps get entities precisely. However, SPARQL is not suitable for most repository users who explore biosimulation models freely without adequate knowledge regarding ontologies, RDF structure, and SPARQL syntax. We propose here a text-based information retrieval approach, CASBERT, that is easy to use and can present candidates of relevant entities from models across a repository's contents. CASBERT adapts Bidirectional Encoder Representations from Transformers (BERT), where each composite annotation about an entity is converted into an entity embedding for subsequent storage in a list-like structure. For entity lookup, a query is transformed to a query embedding and compared to the entity embeddings, and then the entities are displayed in order based on their similarity. The simple list-like structure makes it possible to implement CASBERT as an efficient search engine product, with inexpensive addition, modification, and insertion of entity embedding. To demonstrate and test CASBERT, we created a dataset for testing from the Physiome Model Repository and a static export of the BioModels database consisting of query-entities pairs. Measured using Mean Average Precision and Mean Reciprocal Rank, we found that our approach can perform better than the traditional bag-of-words method.

show abstract

Section: Methodsmentioning

confidence: 99%

CASBERT: BERT-Based Retrieval for Compositely Annotated Biosimulation Model Entities

Munarko

Rampadarath

Nickerson

2022

Preprint

View full text Add to dashboard Cite

show abstract

“…Glove [23]). A recent study [24] empirically shows that classical pretrained embeddings can match contextual embeddings on industry-scale data, and often perform within 5 to 10% accuracy (absolute) on benchmark tasks.…”

Section: J O U R N a L P R E -P R O O Fmentioning

confidence: 99%

Accurate Clinical and Biomedical Named Entity Recognition at Scale

Kocaman

Talby

2022

Software Impacts

View full text Add to dashboard Cite

“…Compared with contextualized embeddings, static embeddings like Skipgram (Mikolov et al, 2013) and GloVe (Pennington et al, 2014) are lighter and less computationally expensive. Furthermore, they can even perform without significant performance loss for contextindependent tasks like lexical-semantic tasks (e.g., word analogy), or some tasks with plentiful labeled data and simple language (Arora et al, 2020).…”

Section: Introductionmentioning

confidence: 99%

Using Context-to-Vector with Graph Retrofitting to Improve Word Embeddings

Zheng¹,

Wang²,

Wang³

et al. 2022

Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)

View full text Add to dashboard Cite

Although contextualized embeddings generated from large-scale pre-trained models perform well in many tasks, traditional static embeddings (e.g., Skip-gram, Word2Vec) still play an important role in low-resource and lightweight settings due to their low computational cost, ease of deployment, and stability. In this paper, we aim to improve word embeddings by 1) incorporating more contextual information from existing pre-trained models into the Skip-gram framework, which we call Context-to-Vec; 2) proposing a post-processing retrofitting method for static embeddings independent of training by employing priori synonym knowledge and weighted vector distribution. Through extrinsic and intrinsic tasks, our methods are well proven to outperform the baselines by a large margin.

show abstract

Contextual Embeddings: When Are They Worth It?

Cited by 8 publications

References 7 publications

CASBERT: BERT-Based Retrieval for Compositely Annotated Biosimulation Model Entities

CASBERT: BERT-Based Retrieval for Compositely Annotated Biosimulation Model Entities

Accurate Clinical and Biomedical Named Entity Recognition at Scale

Using Context-to-Vector with Graph Retrofitting to Improve Word Embeddings

Contact Info

Product

Resources

About