Repository navigation
Conversation
…lock ParseSitemaps returned from the goroutine without calling wg.Done() when sc.client.Get failed, leaving the WaitGroup counter permanently above zero. Crawler.Start blocks on wg.Wait() before reaching the crawl loop, so the crawl never starts and never times out: the context deadline is only checked inside crawl(). The project stays in "Crawling" state forever with zero crawled URLs.
Contributor
Author
|
Sorry, I didn't notice that such a PR already exists. |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Problem
A crawl can hang forever before a single URL is fetched. The project stays in
the "Crawling" state indefinitely, with 0 crawled URLs, and never recovers.
I hit this on three WordPress sites whose Yoast sitemap index takes longer than
the 10s
ClientTimeoutto generate. Goroutine dump after ~48 hours:Root cause
In
ParseSitemaps,wg.Done()is the last statement of the goroutine, so theearly
returnon a failed request skips it:The counter stays above zero and
wg.Wait()blocks forever.The 2h
context.WithTimeoutinNewCrawlerdoes not help.Start()callsParseSitemapsbefore the crawl loop, and the context is only observed insidecrawl(), which is never reached:Second issue
checkIndexcallssitemap.ParseIndexFromSite, which fetches throughhttp.Getonhttp.DefaultClient- no timeout at all. An unresponsive sitemapindex hangs the crawl with no stack trace pointing anywhere useful. That request
also bypassed the project's User-Agent and basic auth settings.
Changes
defer wg.Done()as the first statement of the goroutine inParseSitemaps.checkIndexfetches the index throughsc.clientand parses it withsitemap.ParseIndex, so the configured timeout, User-Agent and basic authapply.
ParseSitemapsreturns instead of blocking.
Reproducing
Any site whose sitemap takes longer than
ClientTimeout(10s) to respond, with"Crawl sitemap" enabled. A large WordPress install with an uncached Yoast
sitemap index is the common case. The new test reproduces it without network
access - on
mainit hangs until the test timeout.Notes
The two fixes are independent and are in separate commits, in case you want only
the first one.
ClientTimeoutis left at 10s. After this change a slow sitemap is skippedrather than crawled, which may deserve a separate discussion - a longer timeout
just for sitemap requests would be reasonable, since they are far larger than a
typical page.