,Implementation of a weblog extraction system with an improved template extraction technique

来源 :中国文献情报(英文刊) | 被引量 : 0次 | 上传用户:pgwork2011
下载到本地 , 更方便阅读
声明 : 本文档内容版权归属内容提供方 , 如果您对本文有版权争议 , 可与客服联系进行内容授权或下架
论文部分内容阅读
Purpose:The objectives of this study are to explore an effective technique to extract information from weblogs and develop an experimental system to extract structured information as much as possible with this technique.The system will lay a foundation for evaluation,analysis,retrieval,and utilization of the extracted information.Design/methodology/approach:An improved template extraction technique was proposed.Separate templates designed for extracting blog entry titles,posts and their comments were established,and structured information was extracted online step by step.A dozen of data items,such as the entry titles,posts and their commenters and comments,the numbers of views,and the numbers of citations were extracted from eight major Chinese blog websites,including Sina,Sohu and Bokee.Findings:Results showed that the average accuracy of the experimental extraction system reached 94.6%.Because the online and multi-threading extraction technique was adopted,the speed of extraction was improved with the average speed of 15 pages per second without considering the network delay.In addition,entries posted by Ajax technology can be extracted successfully.Research limitations:As the templates need to be established in advance,this extraction technique can be effectively applied to a limited range of blog websites.In addition,the stability of the extraction templates was affected by the source code of the blog pages.Practical implications:This paper has studied and established a blog page extraction system,which can be used to extract structured data,preserve and update the data,and facilitate the collection,study and utilization of the blog resources,especially academic blog resources.Originality/value:This modified template extraction technique outperforms the Web page downloaders and the specialized blog page downloaders with structured and comprehensive data extraction.
其他文献
在新形势下,加强和改进思想政治工作,增强思想政治工作的实效性,创造性提出了八个结合元素--“八个结合八个反对”,即继承与创新、物质鼓励和精神鼓励、情与理、言教与身教、
Purpose:This research looks at the characteristics of social networks of rural people in different careers and income conditions.The study also investigates inf
信息网络的发展使高校思想政治工作的教育机制、教育职能、施教途径以及教师职业角色等方面面临着更多的挑战.当前高校思想政治工作网络化要坚持主动性、主导性、主体性的原
大学生是构建社会主义和谐社会的重要力量,做好大学生的思想政治教育工作,具有长远的战略意义.因此,必须结合和谐社会的时代特征和当代大学生的思想现状,加强和改进大学生的
In relation to the service strategies,service systems,and service items provided by CALIS,this paper introduced the service architecture of China Academic Libra
毕业生科学择业观的树立是缓解就业压力长期存在的需要,也是帮助社会合理配置人才资源,实施科教兴国战略的重要举措.在大学毕业生中加强思想政治教育,是引导毕业生树立科学的
成人教育担负着在职从业人员和求职待业人员等的教育培训任务,提高成人教育对象的思想道德素质,促进他们的全面发展,是成人教育德育极其重要的职责。在构建和谐社会的背景下,
Evaluating government openness is important in monitoring government performance and promoting government transparency. Therefore, it is necessary to develop an
构建和谐社会和大学生思想政治教育两者之间有着相辅相成的关系,加强和改进大学生思想政治教育,有效地为构建和谐社会服务.在当前复杂的社会环境下,大学生在思想政治上、在理
随着多媒体应用的普及,图像压缩的研究变得越来越重要。本文通过对游程编码与JPEG压缩编码标准的介绍,着重讨论了JPEG的压缩原理。实验证明,JPEG编码的确具有很高的压缩比,所