Skip to content

Fix crawl hanging when a sitemap request fails - #52

Merged
StJudeWasHere merged 1 commit into
StJudeWasHere:mainfrom
davidpelayo:fix/sitemap-waitgroup-deadlock
Oct 8, 2026
Merged

StJudeWasHere merged 1 commit into
StJudeWasHere:mainfrom
davidpelayo:fix/sitemap-waitgroup-deadlock

Conversation

@davidpelayo

Copy link
Copy Markdown
Contributor

Fixes #51.

Problem

SitemapChecker.ParseSitemaps calls wg.Add(1) for every sitemap, but wg.Done() was only reached on the success path. When sc.client.Get returned an error the goroutine returned early, the WaitGroup counter never reached zero, and wg.Wait() blocked forever.

ParseSitemaps is called from Crawler.Start, so a single failed sitemap request hung the entire crawl before a single page was fetched — the UI sits on "Crawling..." indefinitely with no page reports and no outbound requests.

A transient error is enough to hang the crawl permanently. Sites whose sitemap index fans out to many children are more exposed, because the children are requested concurrently: the more parallel requests, the higher the chance that one fails.

Fix

Defer wg.Done() so the WaitGroup is released on every path, including the early return. Unreachable sitemaps are skipped and the remaining ones are still parsed and crawled.

 go func(s string) {
+	defer wg.Done()
+
 	resp, err := sc.client.Get(s)
 	if err != nil {
 		return
 	}
 	...
-	wg.Done()
 }(s)

Tests

sitemap_checker.go had no test file. This adds internal/crawler/sitemap_checker_test.go covering:

  • TestParseSitemapsReturnsWhenRequestFails — every sitemap request fails; ParseSitemaps must still return.
  • TestParseSitemapsWithFailingAndWorkingSitemaps — one unreachable sitemap must not discard the working ones. This is the real-world case: an index with several children where one request fails.
  • TestParseSitemapsRespectsLimit — the crawl limit is still honoured.

Two details worth noting:

  • The tests run ParseSitemaps in a goroutine guarded by a timeout, so a regression fails the test rather than hanging the whole test run. Without that, a reintroduced bug would stall CI instead of reporting.
  • They use a local httptest server for the sitemap index, because ParseSitemaps resolves the index over the network via checkIndex. This keeps the tests hermetic and fast.

Verified against the exact CI commands:

$ go test -race -parallel 1 -short ./internal/...
ok  	github.com/stjudewashere/seonaut/internal/config	1.629s
ok  	github.com/stjudewashere/seonaut/internal/crawler	2.029s
ok  	github.com/stjudewashere/seonaut/internal/issues/page	2.590s
ok  	github.com/stjudewashere/seonaut/internal/repository	2.699s
ok  	github.com/stjudewashere/seonaut/internal/services	15.287s
ok  	github.com/stjudewashere/seonaut/internal/urlutils	1.910s

$ go build -race ./cmd/server/main.go

Confirmed the tests actually catch the bug — reverting sitemap_checker.go to main and rerunning:

--- FAIL: TestParseSitemapsReturnsWhenRequestFails (5.00s)
    sitemap_checker_test.go:124: ParseSitemaps did not return within 5s; the WaitGroup was never released
--- FAIL: TestParseSitemapsWithFailingAndWorkingSitemaps (5.00s)
    sitemap_checker_test.go:138: ParseSitemaps did not return within 5s; the WaitGroup was never released
FAIL

gofmt -l and go vet ./internal/crawler/ are both clean.

Notes

Beyond the scope of this PR, but noticed while debugging: a crawl that stalls has no timeout or watchdog, so any future hang in Crawler.Start presents the same way — an indefinite "Crawling..." with no diagnostics. A crawl-level timeout, or logging failed sitemap requests instead of silently swallowing the error, would make this class of problem much easier to spot. Happy to open a separate issue if that is of interest.

wg.Add(1) was called for every sitemap, but wg.Done() was only reached on
the success path. When sc.client.Get returned an error the goroutine
returned early, leaving the WaitGroup counter above zero, so wg.Wait()
blocked forever. ParseSitemaps is called from Crawler.Start, so a single
failed sitemap request hung the whole crawl before any page was fetched.

A transient error is enough to hang the crawl permanently. Sites whose
sitemap index fans out to many children are more exposed, since the
children are requested concurrently.

Deferring wg.Done() releases the WaitGroup on every path. Unreachable
sitemaps are skipped and the remaining ones are still parsed.

Adds tests for ParseSitemaps, which had no test file:

- returns when every sitemap request fails
- one failing sitemap does not discard the working ones
- the crawl limit is respected

The tests run ParseSitemaps in a goroutine with a timeout so a regression
fails the test instead of hanging the test run. They use a local httptest
server for the sitemap index, since ParseSitemaps resolves the index over
the network.
@davidpelayo
davidpelayo force-pushed the fix/sitemap-waitgroup-deadlock branch from dc1343d to 91f67d7 Compare July 31, 2026 13:47
@davidpelayo davidpelayo mentioned this pull request Jul 31, 2026
@StJudeWasHere
StJudeWasHere merged commit cc97e1e into StJudeWasHere:main Oct 8, 2026
@StJudeWasHere

Copy link
Copy Markdown
Owner

Thanks for the fix and the detailed PR.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Crawl hangs forever when a sitemap request fails (WaitGroup never released in ParseSitemaps)

2 participants