面向Web论坛的网络信息获取技术及系统实现

doi:10.3969/j.issn.1007130X.2011.

J4 ›› 2011, Vol. 33 ›› Issue (1): 157-160.doi: 10.3969/j.issn.1007130X.2011.

面向Web论坛的网络信息获取技术及系统实现

彭冬，蔡皖东

(西北工业大学计算机学院，陕西西安 710072)

收稿日期:2010-01-04 修回日期:2010-05-13 出版日期:2011-01-25 发布日期:2011-01-25
通讯作者: 彭冬 E-mail:justme@mail.nwpu.edu.cn
作者简介:彭冬（1984），男，四川江油人，硕士,研究方向为网络与信息安全。蔡皖东（1955），男,山东文登人，博士，教授，研究方向为网络与信息安全。
基金资助:
国家863计划资助项目（2009AA01Z424）;2009届西北工业大学本科毕业设计重点扶持项目

The Web Forum Crawling Technology and System Implementation

PENG Dong,CAI Wandong

(School of Computer Science,Northwestern Polytechnical University,Xi’an 710072,China)

Received:2010-01-04 Revised:2010-05-13 Online:2011-01-25 Published:2011-01-25

摘要/Abstract

摘要：

网络爬虫技术是网络信息获取的重要手段，面向Web论坛的信息获取则是网络爬虫技术所面临的新课题。在分析和研究面向Web论坛信息获取技术的基础上，本文设计和实现了一种用于Web论坛信息获取的主题网络爬虫系统，根据Web论坛信息组织结构，提出了基于遍历策略的信息搜索技术；根据正文信息分布及论坛自身特点，提出了基于DOM与分块算法相结合的正文提取技术。实验结果表明，遍历策略比传统的网络爬虫遍历策略具有更高的效率，能够采集到更多主题相关度高的网页；经过噪声清洗处理后，有效提取网页正文，提高了信息采集精度。

关键词: 网络爬虫, Web论坛, 正文提取, 主题相关度

Abstract:

The Web spider is very important in gathering information, which also faces new challenges when it's been used in crawling the Web forum. This paper mainly studies the basic technologies of crawling in the Web forum, designs and implements such a system, which is mainly used to gather the information of the Web forum. According to the information structure, a traversal strategy is proposed. Based on the distribution of the context, a DOM and block algorithm is proposed. The experimental result shows that the traversal strategy is more efficient than the traditional traverses to get those highly subjectrelevant Web pages, and after using the strategy for the context extracting of Web pages, effectively improves the accuracy of the information collection.

Key words: web spider;web forum;context extracting;subject relevant

彭冬，蔡皖东. 面向Web论坛的网络信息获取技术及系统实现[J]. J4, 2011, 33(1): 157-160.

PENG Dong,CAI Wandong. The Web Forum Crawling Technology and System Implementation[J]. J4, 2011, 33(1): 157-160.

[1]	于娟，刘强. 主题网络爬虫研究综述[J]. J4, 2015, 37(02): 231-237.
[2]	屈振新，朱文昌. 基于云计算的定向搜索监控研究[J]. J4, 2013, 35(1): 82-87.
[3]	王振宇1，唐远华1，郭力2. 面向分层结构的网页分类与抓取[J]. J4, 2012, 34(11): 1-6.
[4]	范会联1，李献礼2，曾广朴1. 基于改进遗传算法的聚焦爬虫设计[J]. J4, 2010, 32(5): 126-129.
[5]	程菲汪建海罗键. 增量更新Crawler进行Web收集方法研究[J]. J4, 2006, 28(12): 28-30.

面向Web论坛的网络信息获取技术及系统实现

The Web Forum Crawling Technology and System Implementation

PDF

可视化

摘要/Abstract

引用本文

使用本文

相关文章 5

编辑推荐

Metrics

本文评价