基于多种策略的页面内容提取算法 Web Content Extraction Based on Multiple Strategies期刊界 All Journals 搜尽天下杂志传播学术成果专业期刊搜索期刊信息化学术搜索

基于多种策略的页面内容提取算法

引用本文：	高琰,谷士文,谭立球. 基于多种策略的页面内容提取算法[J]. 西南交通大学学报, 2007, 42(4): 473-477

作者姓名：	高琰谷士文谭立球

作者单位：	中南大学信息科学与工程学院,湖南,长沙,410075

摘要：	针对W eb页面存在与主题无关的噪音的问题,提出了基于页面结构与页面内容相结合的多策略页面内容提取算法.该算法根据改进的VIPS(基于视觉信息的页面分割算法)生成页面的块结构树,通过定义内聚度阈值和块结构树的最大深度,实现了块结构树中不同区域内不同分块粒度的要求;根据W eb页面提供的结构信息和内容信息提取块结构树叶子节点中的"主题"块和"主题相关"块;最后,对主题块和主题相关块的内容进行合并,提取页面的主要内容.实验表明,对任意下载、不同内容类型的页面,该算法都能有效地提取页面内容.
关键词：	VIPS(基于视觉信息的页面分割算法) 内聚度最大深度内容信息结构信息
文章编号：	0258-2724（2007）04-0473-05
修稿时间：	2006-06-14
Web Content Extraction Based on Multiple Strategies

GAO Yan,GU Shiwen,TAN Liqiu. Web Content Extraction Based on Multiple Strategies[J]. Journal of Southwest Jiaotong University, 2007, 42(4): 473-477

Authors:	GAO Yan GU Shiwen TAN Liqiu

Affiliation:	College of Information Science and Eng. , Central South University, Changsha 410075, China

Abstract:	In order to filter the noise in a web page,a new multi-strategy algorithm to extract the contents of a web page was proposed.With this algorithm,the granularity in different areas of the block tree of a web page established by the improved VIPS(visual based page segment) algorithm is controlled by defining the permitted degree of coherence and the maximum depth of the block tree.In addition,"topic" or "topic-relevant" blocks among the leaves of the block tree can be extracted from the blocks' content information and structure information.Finally,the main content of a web page can be extracted by merging these blocks' contents.Experiments on the web pages of three sites indicates that the proposed algorithm is effective for extracting the contents of any type of web pages.

Keywords:	VIPS(visual based page segment) degree of coherence maximum depth content information structure information
本文献已被 CNKI 维普万方数据等数据库收录！

设为首页 | 免责声明 | 关于勤云 | 加入收藏