Ramirez-Orta, Juan scite author profile

Ramirez-Orta, Juan

3Publications

2Citation Statements Received

20Citation Statements Given

How they've been cited

How they cite others

Affiliations

CTIC Foundation

Publications

Order By: Most citations

Unsupervised document summarization using pre-trained sentence embeddings and graph centrality

Juan¹,

Milios²

2021

View full text Add to dashboard Cite

This paper describes our submission for the LongSumm task in SDP 2021. We propose a method for incorporating sentence embeddings produced by deep language models into extractive summarization techniques based on graph centrality in an unsupervised manner. The proposed method is simple, fast, can summarize any document of any size and can satisfy any length constraints for the summaries produced. The method offers competitive performance to more sophisticated supervised methods and can serve as a proxy for abstractive summarization techniques.

show abstract

Post-OCR Document Correction with large Ensembles of Character Sequence-to-Sequence Models

Juan¹,

Xamena²,

Maguitman³

et al. 2021

Preprint

View full text Add to dashboard Cite

In this paper, we propose a novel method based on character sequence-to-sequence models to correct documents already processed with Optical Character Recognition (OCR) systems. The main contribution of this paper is a set of strategies to accurately process strings much longer than the ones used to train the sequence model while being sample-and resource-efficient, supported by thorough experimentation. The strategy with the best performance involves splitting the input document in character n-grams and combining their individual corrections into the final output using a voting scheme that is equivalent to an ensemble of a large number of sequence models. We further investigate how to weigh the contributions from each one of the members of this ensemble. We test our method on nine languages of the ICDAR 2019 competition on post-OCR text correction and achieve a new state-of-the-art performance in five of them. Our code for post-OCR correction is shared at (omitted in this draft to enable blind review).

show abstract

Algoritmos fonéticos para la detección de palabras fonéticamente similares en el español del centro de México

Hernández-Mena

Meza

Juan

et al. 2020

RESLA

View full text Add to dashboard Cite

Resumen En la actualidad, la detección de palabras fonéticamente similares se ha logrado de forma exitosa gracias a la utilización de algoritmos fonéticos. Sin embargo, tales algoritmos dependen del lenguaje al que pertenecen, por lo que generalmente no están optimizados para el español. Por esta razón, en el siguiente artículo se presentará el algoritmo PFS y su variante PFS-US, los cuales son algoritmos fonéticos que consideran la fonología del español hablado en el centro de México, y fueron diseñados para detectar palabras fonéticamente similares en grandes conjuntos de palabras. Ahora bien, a través de un análisis comparativo entre otros cuatro algoritmos fonéticos de estado del arte, analizaremos la consideración fonológica mencionada. Para ello, se definieron métricas independientes de la lengua para evaluar algoritmos fonéticos en general. Dichas métricas se basan en la estructura de los grupos de palabras fonéticamente similares entre sí y su relación con palabras que no son similares con ninguna otra. Adicionalmente, los recursos generados se comparten de forma libre para su uso y análisis.

show abstract

scite is a Brooklyn-based organization that helps researchers better discover and understand research articles through Smart Citations–citations that display the context of the citation and describe whether the article provides supporting or contrasting evidence. scite is used by students and researchers from around the world and is funded in part by the National Science Foundation and the National Institute on Drug Abuse of the National Institutes of Health.

Contact Info

hi@scite.ai

10624 S. Eastern Ave., Ste. A-614

Henderson, NV 89052, USA

Blog Terms and Conditions API Terms Privacy Policy Contact Cookie Preferences Do Not Sell or Share My Personal Information

Made with 💙 for researchers

Part of the Research Solutions Family.