Topic Extraction for Documents Based on Compressibility Vector
スポンサーリンク
概要
- 論文の詳細を見る
Nowadays, there are a great deal of e-documents being accessed on the Internet. It would be helpful if those documents and significant extract contents could be automatically analyzed. Similarity analysis and topic extraction are widely used as document relation analysis techniques. Most of the methods being proposed need some processes such as stemming, stop words removal, and etc. In those methods, natural language processing (NLP) technology is necessary and hence they are dependent on the language feature and the dataset. In this study, we propose novel document relation analysis and topic extraction methods based on text compression. Our proposed approaches do not require NLP, and can also automatically evaluate documents. We challenge our proposal with model documents, URCS and Reuters-21578 dataset, for relation analysis and topic extraction. The effectiveness of the proposed methods is shown by the simulations.
- The Institute of Electronics, Information and Communication Engineersの論文
著者
関連論文
- Vibration Characteristics of Cascaded Blades of a Transonic Turbine : Part 1 : Experiment
- Vibration Characteristics of Cascaded Blades of a Transonic Turbine : Part 2 : Numerical Analysis
- Topic Extraction for Documents Based on Compressibility Vector