|
Style |
|
|
|
|
|
High-performance parallel web crawling |
|
UbiCrawler is a scalable, fault-tolerant and fully distributed web crawler developed in collaboration with the Istituto di Informatica e Telematica. The first report on the design of UbiCrawler won the Best Poster Award at the Tenth World Wide Web Conference. |
|
|
Once a part of the web has been crawled, the resulting graph is very large—you need a compact representation. WebGraph is a framework built to this purpose. Among other things, WebGraph uses new instantaneous codes for the integers and new aggressive algorithmic compression techniques. |
|
|
Web graphs have special properties whose study requires a sizeable amount of mathematics, but also
a careful study of actual web graphs. We have studied, for instance, the paradoxical way PageRank evolves during a crawl, and the way PageRank changes depending on the damping factor. |
|
|
Search-engine construction |
|
Often, the purpose of a crawl is the contruction of a full-text index of the text contained in the crawled pages. Such an index is at the basis of all existing commercial search engines such as Google.
The research of search-engine construction is based on MG4J, a system for full-text indexing of large-scale document collections. |
|
| |
|
|
|
|
Newsflash |
|
The new LAW software release (1.3) provides code for reading WARC/0.9 web archive files. |
|
|
|
News |
|
|
|
|