Cooperative strategy for web data mining and cleaning

Li, Yuefeng, Zhang, Chengqi, & Zhang, Shichao (2003) Cooperative strategy for web data mining and cleaning. Applied Artificial Intelligence, 17(5-6), pp. 443-460.

View at publisher


While the Internet and World Wide Web have put a huge volume of low-quality information at the easy access of an information gathering system, filtering out irrelevant information has become a big challenge. In this paper, a Web data mining and cleaning strategy for information gathering is proposed. A data-mining model is presented for the data that come from multiple agents. Using the model, a data-cleaning algorithm is then presented to eliminate irrelevant data. To evaluate the data-cleaning strategy, an interpretation is given for the mining model according to evidence theory. An experiment is also conducted to evaluate the strategy using Web data. The experimental results have shown that the proposed strategy is efficient and promising.

Impact and interest:

18 citations in Scopus
15 citations in Web of Science®
Search Google Scholar™

Citation counts are sourced monthly from Scopus and Web of Science® citation databases.

These databases contain citations from different subsets of available publications and different time periods and thus the citation count from each is usually different. Some works are not in either database and no count is displayed. Scopus includes citations from articles published in 1996 onwards, and Web of Science® generally from 1980 onwards.

Citations counts from the Google Scholar™ indexing service can be viewed at the linked Google Scholar™ search.

Full-text downloads:

290 since deposited on 19 Apr 2007
8 in the past twelve months

Full-text downloads displays the total number of times this work’s files (e.g., a PDF) have been downloaded from QUT ePrints as well as the number of downloads in the previous 365 days. The count includes downloads for all files if a work has more than one.

ID Code: 7050
Item Type: Journal Article
Refereed: Yes
Keywords: Web mining, information fusion, information agents
DOI: 10.1080/713827173
ISSN: 1087-6545
Subjects: Australian and New Zealand Standard Research Classification > INFORMATION AND COMPUTING SCIENCES (080000) > LIBRARY AND INFORMATION STUDIES (080700) > Information Retrieval and Web Search (080704)
Australian and New Zealand Standard Research Classification > INFORMATION AND COMPUTING SCIENCES (080000) > ARTIFICIAL INTELLIGENCE AND IMAGE PROCESSING (080100)
Divisions: Past > QUT Faculties & Divisions > Faculty of Science and Technology
Copyright Owner: Copyright 2003 Taylor & Francis
Copyright Statement: First published in Applied Artificial Intelligence 17(5-6):pp. 443-460.
Deposited On: 19 Apr 2007 00:00
Last Modified: 22 Apr 2013 03:54

Export: EndNote | Dublin Core | BibTeX

Repository Staff Only: item control page