Creating a web crawler

I am currently developing a custom search engine with a built-in crawler. For some reason I am not multithreaded, so so far my index has been single-threaded encoded. Now I have a little dilemma with the caterpillar I am building. Can anyone suggest which is better, crawl 1 page then index it, or crawl 1000+ page and cache and then index it?

+1


a source to share


4 answers


Networks are slow (relative to the CPU). You will see a significant increase in speed by parallelizing your finder. Otherwise, your application will spend most of its time waiting for network I / O to complete. You can use multiple threads and block IO or one thread with asynchronous IO.



In addition, most indexing algorithms will perform better on batches of documents that index one document at a time.

+4


a source


better? From the point of view of what? In terms of speed, I can't see a noticeable difference. From a reliability standpoint (recovering from a catastrophic failure), it is probably best to index every page as it crawls.



+1


a source


I would strongly suggest getting "in" multithreading if you are serious about your finder. Basically, you would like to have at least one indexer and at least one crawler (potentially many for both) executed at any time. Among other things, it minimizes startup and shutdown costs (for example, initializing and freeing data structures).

+1


a source


Not using streams is ok. However, if you still want performance, you need to deal with asynchronous IO. I would recommend checking the Boost.ASIO link text . Using asynchronous I / O will make your dilemma "irrelevant" because it doesn't matter. Also as a bonus, in the future, if you decide to use threads, then its trivial to tell Boost.Asio to apply multithreading to the problem.

+1


a source







All Articles