The "Product Bible" — features, user stories, design system, and brand guidelines. Generated by the Torvaldsen workflow. This is the canonical source for product requirements.
Parsify is a documentation-as-a-service API that transforms web documentation into AI-ready formats. Built on top of the proven DocScraper toolkit, Parsify exposes scraping, cleaning, chunking, and classification capabilities through a simple REST API.
The target audience is development teams building RAG (Retrieval-Augmented Generation) pipelines, knowledge bases, and AI-powered search systems. These teams currently run custom scraping scripts, manually clean content, and struggle with maintaining documentation freshness. Parsify eliminates this operational burden by providing a managed, reliable, and scalable document processing service.
The unique value proposition is the combination of intelligent cleaning (25+ content patterns), AI-powered classification (GPT-4o), semantic chunking optimized for embeddings, and dependency analysis between documents — all accessible via API.
Purpose: Establish a stable, correctly-tested, well-documented codebase before building any API layer.
User-Facing Description: Internal quality milestone — users won't interact with this directly.
Technical Requirements:
- Fix
output_dirparameter being silently ignored inDocPostProcessor.process_documents()(line 622) - Replace all bare
except:clauses with specific exception types - Parameterize all hardcoded file paths (especially
/Volumes/NvME-Satechi/references) - Use
pathlib.Pathinstead ofos.paththroughout - Remove
logging.basicConfig()from library modules (only in entry points) - Add type hints to all public methods
- Resolve
DocumentationScraperclass name collision betweenDocScraper.pyandSimpleDocScraper.py - Fix LLM cost double-counting in
PostScraperCleaner.py
Data Requirements:
- No data model changes — this is code quality work
Edge Cases & Constraints:
- GUI files use bare excepts for Tkinter event handling — may need careful exception selection
- Some test assertions depend on exact pattern counts which change as patterns are added
Acceptance Criteria:
-
python -m pytest tests/ -vpasses with all tests green - No hardcoded
/Volumes/paths in any source file - No bare
except:clauses in any source file - No
logging.basicConfig()calls in library modules (only in__main__blocks or scripts) - All public methods have type hints
- ONBOARDING.json factual errors corrected
Purpose: Enable programmatic access to documentation scraping and processing capabilities.
User-Facing Description: Submit a URL and get back cleaned, chunked, classified documentation through a simple REST API with API key authentication.
Technical Requirements:
- FastAPI application with versioned routes (
/api/v1/) - Pydantic models for all request/response schemas
- Async job queue (ARQ + Redis) for scraping operations (5-30 min runtime)
- API key authentication (SHA-256 hashed, prefix-based lookup)
- Job status polling and webhook callbacks
- Result download (zip archive of processed documents)
- Input validation: URL format, allowed domains (optional), max depth, max pages
Data Requirements:
Jobmodel: id, user_id, status (queued/running/completed/failed), url, config, result_path, created_at, completed_at, error_messageAPIKeymodel: id, user_id, key_prefix, key_hash, name, created_at, last_used_at, revokedUsermodel: id, email, name, created_at, subscription_tier
Edge Cases & Constraints:
- Scraping jobs may take 30+ minutes — must handle timeouts gracefully
- Target sites may block scrapers — return clear error with retry guidance
- Large documentation sets may produce >1GB of output — implement size limits
- Rate limiting must work per-key, not per-IP (for shared infrastructure)
- Does NOT include billing (that's F002) or persistent storage (that's F003)
Acceptance Criteria:
- POST /api/v1/scrape accepts URL + config, returns job ID
- GET /api/v1/jobs/{id} returns job status with progress percentage
- GET /api/v1/jobs/{id}/result returns downloadable results when complete
- POST /api/v1/process accepts uploaded documents, returns processed output
- Invalid API key returns 401 with clear error message
- Expired/revoked key returns 401
- Job queue handles concurrent scraping jobs without interference
- Webhook fires on job completion/failure (if URL provided)
- OpenAPI spec auto-generated at /docs
Purpose: Enable sustainable business model with self-service billing.
User-Facing Description: Monitor your API usage, manage your subscription, and pay for what you use through an integrated billing system.
Technical Requirements:
- Per-key rate limiting (requests/minute and pages/month)
- Usage metering: API calls, pages scraped, AI tokens consumed
- Stripe integration: checkout, subscription management, webhooks
- Tier enforcement: free (100 pages/mo), pro (10K pages/mo), enterprise (custom)
- Usage dashboard: current period usage, historical charts
- Overage handling: soft limit warnings, hard limit enforcement
Data Requirements:
UsageRecordmodel: id, user_id, api_key_id, endpoint, pages_scraped, tokens_used, timestampSubscriptionmodel: id, user_id, stripe_subscription_id, tier, status, current_period_start, current_period_end
Edge Cases & Constraints:
- Usage tracking must be atomic (no double-counting on retries)
- Stripe webhooks must be idempotent (handle duplicate events)
- Downgrade mid-cycle: prorate and adjust limits immediately
- Free tier must work without Stripe (no credit card required)
- Does NOT include admin dashboard (Phase 4+) or custom enterprise contracts
Acceptance Criteria:
- Rate limiter returns 429 with retry-after header when exceeded
- Usage counter accurately tracks pages scraped per API key
- Free tier enforced at 100 pages/month without credit card
- Pro tier checkout creates Stripe subscription
- Subscription webhook updates user tier in database
- Usage dashboard endpoint returns current period stats
Purpose: Make the service production-ready with proper database, CI/CD, and deployment.
User-Facing Description: Reliable, always-available service with automated deployments and monitoring.
Technical Requirements:
- PostgreSQL schema with Alembic migrations
- Docker multi-stage build (slim runtime image)
- docker-compose for local development (API + PostgreSQL + Redis)
- GitHub Actions: lint (ruff), type-check (mypy), test (pytest), build, deploy
- Production deployment on Railway (or Render)
- Health check endpoint (/health) and readiness probe
- Structured JSON logging with request correlation IDs
- Environment-based configuration (development/staging/production)
Data Requirements:
- Full PostgreSQL schema implementing all models from F001 and F002
- Migration history maintained with Alembic
Edge Cases & Constraints:
- Migration must handle existing data (if any beta users exist)
- Docker image must include Playwright + Chromium (large, ~500MB)
- Redis must be persistent for job queue state
- Does NOT include horizontal scaling (single instance initially)
Acceptance Criteria:
-
docker-compose upstarts full local development stack -
alembic upgrade headapplies all migrations cleanly - GitHub Actions CI passes on every PR
- Production deployment triggered on merge to main
- Health check returns 200 with service status
- Structured logs include request_id for tracing
Purpose: Enable developer discovery, evaluation, and adoption.
User-Facing Description: Comprehensive API documentation with examples, a landing page explaining what Parsify does, and integration guides.
Technical Requirements:
- MkDocs documentation site with Material theme
- Auto-generated API reference from OpenAPI spec
- Getting started guide with curl, Python, and JavaScript examples
- Authentication guide
- Webhook integration guide
- Landing page with feature overview and pricing
- SEO-optimized for "documentation scraping API" keywords
Data Requirements:
- None (static content)
Edge Cases & Constraints:
- Documentation must stay in sync with API changes
- Does NOT include blog, changelog (beyond release notes), or community forum
Acceptance Criteria:
- Documentation site deployed and accessible
- API reference matches current endpoints
- Getting started guide works end-to-end (copy-paste and run)
- Landing page loads in <3s
- Pricing page shows all tiers with feature comparison
Purpose: Reduce integration friction with native language SDKs.
User-Facing Description: Install the Parsify SDK in your language and start scraping documentation in 3 lines of code.
Technical Requirements:
- Python SDK auto-generated from OpenAPI (openapi-generator or custom)
- JavaScript/TypeScript SDK auto-generated from OpenAPI
- Published to PyPI and npm respectively
- SDK includes: authentication, job submission, polling, result download
- Async/await support in both SDKs
- Type annotations (Python) and TypeScript definitions (JS)
Data Requirements:
- None (SDK is client-side only)
Edge Cases & Constraints:
- SDK version must be pinned to API version
- Breaking API changes require SDK major version bump
- Does NOT include SDKs for other languages (Go, Ruby, etc.)
Acceptance Criteria:
-
pip install parsifyinstalls Python SDK -
npm install parsifyinstalls JS SDK - SDK examples in documentation work out of the box
- Both SDKs have >90% API coverage
As a backend developer building a RAG pipeline, I want to submit a documentation URL and get back chunked, classified markdown, So that I can feed it directly into my vector database without manual processing.
Journey:
- Developer signs up and gets a free-tier API key
- Submits POST /api/v1/scrape with
{"url": "https://docs.example.com", "max_depth": 3} - Receives job ID and polls status every 30 seconds
- Job completes — downloads zip with cleaned markdown chunks + metadata JSON
- Feeds chunks into Pinecone/Weaviate with the provided metadata
Happy Path: Job completes in 5-10 minutes, produces well-structured chunks Error Path: Target site blocks scraper → returns clear error with suggestion to try different user-agent or contact support Edge Case: Documentation has 10,000+ pages → job respects max_pages limit, returns partial results with continuation token
As a DevOps engineer, I want to schedule weekly re-scrapes of our vendor documentation, So that our knowledge base stays current without manual intervention.
Journey:
- Sets up a cron job that calls POST /api/v1/scrape weekly
- Registers a webhook URL for job completion
- Webhook fires → triggers vector DB re-indexing pipeline
- Monitors usage dashboard to ensure within tier limits
Happy Path: Weekly scrape detects changes, only re-processes modified pages Error Path: Webhook delivery fails → implements retry with exponential backoff Edge Case: Vendor site restructures URLs → scraper handles redirects, logs new URL mappings
As a startup CTO evaluating documentation processing tools, I want to test Parsify's quality on our documentation sources, So that I can decide whether to adopt it for our AI product.
Journey:
- Visits landing page, reads feature overview
- Signs up for free tier (no credit card)
- Runs 3-4 test scrapes on different documentation sites
- Evaluates output quality (chunking, classification, metadata)
- If satisfied, upgrades to pro tier via Stripe checkout
- Distributes API keys to team members
Happy Path: Free tier's 100 pages is enough for evaluation Error Path: Output quality is poor for specific site → provides feedback via support Edge Case: Team needs 50+ API keys → enterprise tier discussion
As a data engineer with existing markdown documentation, I want to upload files for cleaning and chunking without scraping, So that I can standardize my existing documentation for AI consumption.
Journey:
- Uploads a zip of markdown files to POST /api/v1/process
- Specifies processing options (chunk size, classification model, cleaning patterns)
- Receives processed output immediately (synchronous for small batches) or via job (async for large batches)
- Downloads results with metadata and vector-ready chunks
Happy Path: 50 files processed in under 2 minutes Error Path: Uploaded file has encoding issues → returns error with specific file and line Edge Case: Upload exceeds 100MB → returns 413 with guidance to split
Since Parsify is primarily an API service, the "design system" applies to:
-
API response format consistency
-
Documentation site styling
-
Landing page design
-
Developer-First — Every interface decision optimizes for developer experience (clear errors, consistent schemas, copy-paste examples)
-
Predictable — Same patterns everywhere: consistent error formats, pagination, filtering
-
Transparent — Usage, costs, and job status are always visible and accurate
{
"data": { ... },
"meta": {
"request_id": "req_abc123",
"timestamp": "2026-03-22T10:30:00Z"
}
}Error format:
{
"error": {
"code": "RATE_LIMIT_EXCEEDED",
"message": "Rate limit exceeded. Try again in 30 seconds.",
"details": {
"limit": 60,
"remaining": 0,
"reset_at": "2026-03-22T10:31:00Z"
}
},
"meta": {
"request_id": "req_abc123"
}
}| Role | Color | Hex | Usage |
|---|---|---|---|
| Primary | Deep Blue | #1e40af |
CTAs, links, active states |
| Primary Light | Sky Blue | #3b82f6 |
Hover states, highlights |
| Background | White | #ffffff |
Page background |
| Surface | Light Gray | #f8fafc |
Code blocks, cards |
| Text Primary | Slate 900 | #0f172a |
Body text |
| Text Secondary | Slate 500 | #64748b |
Muted text, labels |
| Success | Green | #16a34a |
Success states, "completed" |
| Warning | Amber | #d97706 |
Warnings, "running" |
| Error | Red | #dc2626 |
Errors, "failed" |
| Code | Mono | — | JetBrains Mono for code blocks |
| Element | Font | Weight | Size |
|---|---|---|---|
| H1 | Inter | 700 | 36px |
| H2 | Inter | 600 | 28px |
| H3 | Inter | 600 | 22px |
| Body | Inter | 400 | 16px |
| Code | JetBrains Mono | 400 | 14px |
| API Endpoint | JetBrains Mono | 600 | 16px |
Technical, direct, and reliable. Like Stripe's documentation: precise without being sterile, helpful without being condescending.
| Context | Tone | Example |
|---|---|---|
| Success | Brief, factual | "Job completed. 247 pages processed in 8m 32s." |
| Error | Helpful, actionable | "The target site returned 403 Forbidden. Try setting a custom user-agent header." |
| Empty State | Guiding | "No jobs yet. Submit your first scrape request to get started." |
| Documentation | Clear, example-heavy | Show the curl command first, explain second. |
- Do: Lead with code examples
- Do: Use precise technical language
- Do: State limits explicitly (not "generous free tier" — say "100 pages/month")
- Don't: Use marketing superlatives ("blazing fast", "revolutionary")
- Don't: Hide error details behind generic messages
- Demographics: 3-10 years experience, works at startup or mid-size company
- Goals: Build and maintain AI-powered search/documentation systems
- Pain Points: Manual scraping scripts break, cleaning is tedious, no good managed service
- Communication Preference: Show the API, show the output, let them evaluate
- Demographics: 5+ years experience, manages infrastructure
- Goals: Automate documentation freshness, integrate into CI/CD
- Pain Points: Custom scraping infrastructure is fragile, monitoring is manual
- Communication Preference: Docker compose, helm charts, environment variables
| Metric | Target | How to Measure |
|---|---|---|
| Developer Sign-ups | 100 in first month | User registration count |
| Free → Pro Conversion | 10% within 30 days | Stripe subscription events |
| API Uptime | 99.5% | Health check monitoring |
| Job Success Rate | >95% | Jobs completed / total jobs |
| Median Job Duration | <10 min for 100 pages | Job completion timestamps |
| API Response Time (p95) | <200ms (non-scrape) | Request logging |
| Requirement | Applicability | Implementation |
|---|---|---|
| robots.txt | Yes | Respect target site robots.txt by default |
| Rate Limiting | Yes | Self-imposed limits to avoid hammering target sites |
| Data Retention | Yes | Job results deleted after 7 days (configurable) |
| GDPR | Partial | No PII in scraped content; user accounts use email only |
| Terms of Service | Required | Users responsible for legal right to scrape target content |
| Privacy Policy | Required | Minimal data collection: email, usage metrics, payment info (Stripe) |