Expanding paraphrase lexicons by exploiting generalities

Par Conseil national de recherches du Canada

Téléchargement	Voir la version finale : Expanding paraphrase lexicons by exploiting generalities (PDF, 3.9 Mio)
DOI	Trouver le DOI : https://doi.org/10.1145/3160488
Auteur	Rechercher : Fujita, Atsushi; Rechercher : Isabelle, Pierre¹
Affiliation	Conseil national de recherches du Canada. Technologies numériques
Format	Texte, Article
Résumé	Techniques for generating and recognizing paraphrases, i.e., semantically equivalent expressions, play an important role in a wide range of natural language processing tasks. In the last decade, the task of automatic acquisition of subsentential paraphrases, i.e., words and phrases with (approximately) the same meaning, has been drawing much attention in the research community. The core problem is to obtain paraphrases of high quality in large quantity. This article presents a method for tackling this issue by systematically expanding an initial seed lexicon made up of high-quality paraphrases. This involves automatically capturing morpho-semantic and syntactic generalizations within the lexicon and using them to leverage the power of large-scale monolingual data. Given an input set of paraphrases, our method starts by inducing paraphrase patterns that constitute generalizations over corresponding pairs of lexical variants, such as “amending” and “amendment,” in a fully empirical way. It then searches large-scale monolingual data for new paraphrases matching those patterns. The results of our experiments on English, French, and Japanese demonstrate that our method manages to expand seed lexicons by a large multiple. Human evaluation based on paraphrase substitution tests reveals that the automatically acquired paraphrases are also of high quality.
Date de publication	2018-02-05
Maison d’édition	ACM DL
Dans	ACM Transactions on Asian and Low-Resource Language Information Processing 17, nº 2 : 1–36.
Langue	anglais
Publications évaluées par des pairs	Oui
Numéro NPARC	23002813
Exporter la notice	Exporter en format RIS
Signaler une correction	Signaler une correction (s'ouvre dans un nouvel onglet)
Identificateur de l’enregistrement	b2cf9aa4-4513-4c40-b6e3-e64a2243598f
Enregistrement créé	2018-03-02
Enregistrement modifié	2020-05-30

Date de modification :: 2024-07-07