My crawl of about 400 pages finished with 12 timeouts. I want to fetch those 12 pages again without starting the whole crawl over.
How do I retry only failed pages in a crawl?
2 Answers
14
Accepted answer: resume from the checkpoint file. Failed URLs stay in the checkpoint with their error type, so a resumed run only fetches what did not finish.
crawler.crawl(start_url, checkpoint_path="run.json")
4
You can also collect the failed URLs from the errors list and call batch scrape on just those.