Double-pass clustering technique for multilingual document collections

Research output: Contribution to journalArticle

5 Citations (Scopus)

Abstract

It is often necessary to categorize automatically multilingual document sets, in which documents written in a variety of languages are included, into topically homogeneous subsets, such as when applying an automatic summarization system for multilingual news articles. However, there have been few studies on multilingual document clustering to date. In particular, it is not known whether clustering techniques are effective in medium- or large-scale multilingual document sets. For scalability, techniques should be based on dictionary-based translation and a single- or double-pass clustering algorithm. This article reports on experiments of applying multilingual document clustering to medium-scale sets of English, French, German and Italian documents (Reuters news articles). The results show that the double-pass algorithm has a positive effect in the case that each document is translated. On the other hand, the cluster translation strategy in which clusters obtained by applying a clustering algorithm to each language document set are translated has almost no effect. Also, translation disambiguation techniques can improve, but only slightly, the effectiveness of clustering.

Original languageEnglish
Pages (from-to)304-321
Number of pages18
JournalJournal of Information Science
Volume37
Issue number3
DOIs
Publication statusPublished - 2011 Jun 1

    Fingerprint

Keywords

  • document translation
  • multilingual document clustering
  • translation disambiguation

ASJC Scopus subject areas

  • Information Systems
  • Library and Information Sciences

Cite this