How do I retry only failed pages in a crawl?

My crawl of about 400 pages finished with 12 timeouts. I want to fetch those 12 pages again without starting the whole crawl over.

2 Answers

14

Accepted answer: resume from the checkpoint file. Failed URLs stay in the checkpoint with their error type, so a resumed run only fetches what did not finish.

crawler.crawl(start_url, checkpoint_path="run.json")
4

You can also collect the failed URLs from the errors list and call batch scrape on just those.