Jan Rupnik scite author profile

Jan Rupnik

4Publications

60Citation Statements Received

60Citation Statements Given

How they've been cited

How they cite others

170

Affiliations

B2 (Slovenia), Jožef Stefan Institute, BT Group (United Kingdom)

Publications

Order By: Most citations

News Across Languages - Cross-Lingual Document Similarity and Event Tracking

Rupnik¹,

Muhič²,

Leban³

et al. 2016

jair

View full text Add to dashboard Cite

In today's world, we follow news which is distributed globally. Significant events are reported by different sources and in different languages. In this work, we address the problem of tracking of events in a large multilingual stream. Within a recently developed system Event Registry we examine two aspects of this problem: how to compare articles in different languages and how to link collections of articles in different languages which refer to the same event. Taking a multilingual stream and clusters of articles from each language, we compare different cross-lingual document similarity measures based on Wikipedia. This allows us to compute the similarity of any two articles regardless of language. Building on previous work, we show there are methods which scale well and can compute a meaningful similarity between articles from languages with little or no direct overlap in the training data. Using this capability, we then propose an approach to link clusters of articles across languages which represent the same event. We provide an extensive evaluation of the system as a whole, as well as an evaluation of the quality and robustness of the similarity measure and the linking algorithm.

show abstract

Actionable cognitive twins for decision making in manufacturing

Rožanec

Rupnik

et al. 2021

International Journal of Production Research

View full text Add to dashboard Cite

The Role of Hubs in Cross-Lingual Supervised Document Retrieval

Tomašev

Rupnik

Mladenić

2013

View full text Add to dashboard Cite

A Comparison of Relaxations of Multiset Cannonical Correlation Analysis and Applications

Rupnik¹,

Škraba²,

Shawe‐Taylor³

et al. 2013

Preprint

View full text Add to dashboard Cite

Canonical correlation analysis is a statistical technique that is used to find relations between two sets of variables. An important extension in pattern analysis is to consider more than two sets of variables. This problem can be expressed as a quadratically constrained quadratic program (QCQP), commonly referred to Multi-set Canonical Correlation Analysis (MCCA). This is a non-convex problem and so greedy algorithms converge to local optima without any guarantees on global optimality. In this paper, we show that despite being highly structured, finding the optimal solution is NP-Hard. This motivates our relaxation of the QCQP to a semidefinite program (SDP). The SDP is convex, can be solved reasonably efficiently and comes with both absolute and output-sensitive approximation quality. In addition to theoretical guarantees, we do an extensive comparison of the QCQP method and the SDP relaxation on a variety of synthetic and real world data. Finally, we present two useful extensions: we incorporate kernel methods and computing multiple sets of canonical vectors.

show abstract

scite is a Brooklyn-based organization that helps researchers better discover and understand research articles through Smart Citations–citations that display the context of the citation and describe whether the article provides supporting or contrasting evidence. scite is used by students and researchers from around the world and is funded in part by the National Science Foundation and the National Institute on Drug Abuse of the National Institutes of Health.

Contact Info

hi@scite.ai

10624 S. Eastern Ave., Ste. A-614

Henderson, NV 89052, USA

Blog Terms and Conditions API Terms Privacy Policy Contact Cookie Preferences Do Not Sell or Share My Personal Information

Made with 💙 for researchers

Part of the Research Solutions Family.

Jan Rupnik

News Across Languages - Cross-Lingual Document Similarity and Event Tracking

Actionable cognitive twins for decision making in manufacturing

The Role of Hubs in Cross-Lingual Supervised Document Retrieval

A Comparison of Relaxations of Multiset Cannonical Correlation Analysis and Applications

Contact Info

Product

Resources

About