一种改进的统计与后串最大匹配的中文分词算法研究

J4 ›› 2008, Vol. 30 ›› Issue (8): 79-82.

一种改进的统计与后串最大匹配的中文分词算法研究

吴涛张毛迪陈传波

出版日期:2008-08-01 发布日期:2010-05-19

Online:2008-08-01 Published:2010-05-19

摘要/Abstract

摘要：

在比较各种传统分词方法优缺点的基础上，本文提出了一种新的分词算法。它采用改进的双向Markov链统计方法对词库进行更新，再利用基于词典的有穷自动机后串最大匹配算法以及博弈树搜索算法进行分词。实验结果表明，该分词算法在分词准确性、效率以及生词辨识上取得了良好的效果。

关键词: 正向最大前串匹配逆向最大前串匹配统计法有穷自动机

Abstract:

This paper analyzes several traditional methods for the Chinese word segmentation, compares the advantages and disadvantages of these methods, and presents a new segmentation algorithm. The method adopts the improved bidirectional Markov chain statistical method to update the word library, and then uses the Reverse Maximum Match method based on the word library and the GameTree search algorithm to cut the Chinese word strings. The experimental results show this algorithm has got better effect on veracity, efficiency and new word distinguishment.

Key words: forward maximum match, reverse maximum match, statistical method, definite finite automation

吴涛张毛迪陈传波. 一种改进的统计与后串最大匹配的中文分词算法研究[J]. J4, 2008, 30(8): 79-82.

一种改进的统计与后串最大匹配的中文分词算法研究

PDF

可视化

摘要/Abstract

引用本文

使用本文

相关文章 0

编辑推荐

Metrics

本文评价