检索规则说明:AND代表“并且”;OR代表“或者”;NOT代表“不包含”;(注意必须大写,运算符两边需空一格)
检 索 范 例 :范例一: (K=图书馆学 OR K=情报学) AND A=范并思 范例二:J=计算机应用与软件 AND (U=C++ OR U=Basic) NOT M=Visual
作 者:汤佳杰 曹永忠[1] 顾浩 Tang Jiajie;Cao Yongzhong;Gu Hao(College of Information Engineering,Yangzhou University,Yangzhou,Jiangsu 225000,China)
机构地区:[1]扬州大学信息工程学院
出 处:《计算机时代》2020年第1期69-72,共4页Computer Era
基 金:江苏省研究生研究与实践创新计划KYCX18_2366
摘 要:为了简化网页正文抽取操作与提高网页正文抽取的准确性,提出了一种基于文本标点密度连续和的抽取方法(TPDS)。TPDS基于网页中文本标点分布的密度并计算密度的连续和,选取所有文本块中连续和最大的文本块,将其确定为网页最佳文本块并抽取正文内容。从不同的门户网站随机选取的网页作为测试数据集,实验结果表明,TPDS可有效过滤网页噪声信息得到正文内容。该方法在不同网页上具有很好的适用性,抽取性能优于CETR、CETD、CEPR和CETD-TPC算法。In order to simplify the extraction process of web page text and improve the accuracy of web page text extraction, a method based on text punctuation density continuous sum extraction(TPDS) is proposed. TPDS is based on the density of text punctuation distribution in web pages and calculates the continuous sum of density. The continuous and largest text blocks in all text blocks are selected, which are determined as the best text block of the web page and the body content is extracted. The webpage randomly selected from different portals is used as the test data set. The experimental results show that TPDS can effectively filter the webpage noise information to obtain the body content, and the method has good applicability on different webpage, and the extraction performance is better than CETR, CETD, CEPR and CETD-TPC algorithms.
分 类 号:TP391[自动化与计算机技术—计算机应用技术]
正在载入数据...
正在载入数据...
正在载入数据...
正在载入数据...
正在载入数据...
正在载入数据...
正在载入数据...
正在链接到云南高校图书馆文献保障联盟下载...
云南高校图书馆联盟文献共享服务平台 版权所有©
您的IP:216.73.216.200