Skip to content

Latest commit

 

History

History
436 lines (329 loc) · 17.8 KB

File metadata and controls

436 lines (329 loc) · 17.8 KB

Specification — Parsify

The "Product Bible" — features, user stories, design system, and brand guidelines. Generated by the Torvaldsen workflow. This is the canonical source for product requirements.


Vision

Parsify is a documentation-as-a-service API that transforms web documentation into AI-ready formats. Built on top of the proven DocScraper toolkit, Parsify exposes scraping, cleaning, chunking, and classification capabilities through a simple REST API.

The target audience is development teams building RAG (Retrieval-Augmented Generation) pipelines, knowledge bases, and AI-powered search systems. These teams currently run custom scraping scripts, manually clean content, and struggle with maintaining documentation freshness. Parsify eliminates this operational burden by providing a managed, reliable, and scalable document processing service.

The unique value proposition is the combination of intelligent cleaning (25+ content patterns), AI-powered classification (GPT-4o), semantic chunking optimized for embeddings, and dependency analysis between documents — all accessible via API.


Core Features

F000 — Codebase Cleanup (Audit Fixes)

Purpose: Establish a stable, correctly-tested, well-documented codebase before building any API layer.

User-Facing Description: Internal quality milestone — users won't interact with this directly.

Technical Requirements:

  • Fix output_dir parameter being silently ignored in DocPostProcessor.process_documents() (line 622)
  • Replace all bare except: clauses with specific exception types
  • Parameterize all hardcoded file paths (especially /Volumes/NvME-Satechi/ references)
  • Use pathlib.Path instead of os.path throughout
  • Remove logging.basicConfig() from library modules (only in entry points)
  • Add type hints to all public methods
  • Resolve DocumentationScraper class name collision between DocScraper.py and SimpleDocScraper.py
  • Fix LLM cost double-counting in PostScraperCleaner.py

Data Requirements:

  • No data model changes — this is code quality work

Edge Cases & Constraints:

  • GUI files use bare excepts for Tkinter event handling — may need careful exception selection
  • Some test assertions depend on exact pattern counts which change as patterns are added

Acceptance Criteria:

  • python -m pytest tests/ -v passes with all tests green
  • No hardcoded /Volumes/ paths in any source file
  • No bare except: clauses in any source file
  • No logging.basicConfig() calls in library modules (only in __main__ blocks or scripts)
  • All public methods have type hints
  • ONBOARDING.json factual errors corrected

F001 — REST API Layer

Purpose: Enable programmatic access to documentation scraping and processing capabilities.

User-Facing Description: Submit a URL and get back cleaned, chunked, classified documentation through a simple REST API with API key authentication.

Technical Requirements:

  • FastAPI application with versioned routes (/api/v1/)
  • Pydantic models for all request/response schemas
  • Async job queue (ARQ + Redis) for scraping operations (5-30 min runtime)
  • API key authentication (SHA-256 hashed, prefix-based lookup)
  • Job status polling and webhook callbacks
  • Result download (zip archive of processed documents)
  • Input validation: URL format, allowed domains (optional), max depth, max pages

Data Requirements:

  • Job model: id, user_id, status (queued/running/completed/failed), url, config, result_path, created_at, completed_at, error_message
  • APIKey model: id, user_id, key_prefix, key_hash, name, created_at, last_used_at, revoked
  • User model: id, email, name, created_at, subscription_tier

Edge Cases & Constraints:

  • Scraping jobs may take 30+ minutes — must handle timeouts gracefully
  • Target sites may block scrapers — return clear error with retry guidance
  • Large documentation sets may produce >1GB of output — implement size limits
  • Rate limiting must work per-key, not per-IP (for shared infrastructure)
  • Does NOT include billing (that's F002) or persistent storage (that's F003)

Acceptance Criteria:

  • POST /api/v1/scrape accepts URL + config, returns job ID
  • GET /api/v1/jobs/{id} returns job status with progress percentage
  • GET /api/v1/jobs/{id}/result returns downloadable results when complete
  • POST /api/v1/process accepts uploaded documents, returns processed output
  • Invalid API key returns 401 with clear error message
  • Expired/revoked key returns 401
  • Job queue handles concurrent scraping jobs without interference
  • Webhook fires on job completion/failure (if URL provided)
  • OpenAPI spec auto-generated at /docs

F002 — Billing & Usage Management

Purpose: Enable sustainable business model with self-service billing.

User-Facing Description: Monitor your API usage, manage your subscription, and pay for what you use through an integrated billing system.

Technical Requirements:

  • Per-key rate limiting (requests/minute and pages/month)
  • Usage metering: API calls, pages scraped, AI tokens consumed
  • Stripe integration: checkout, subscription management, webhooks
  • Tier enforcement: free (100 pages/mo), pro (10K pages/mo), enterprise (custom)
  • Usage dashboard: current period usage, historical charts
  • Overage handling: soft limit warnings, hard limit enforcement

Data Requirements:

  • UsageRecord model: id, user_id, api_key_id, endpoint, pages_scraped, tokens_used, timestamp
  • Subscription model: id, user_id, stripe_subscription_id, tier, status, current_period_start, current_period_end

Edge Cases & Constraints:

  • Usage tracking must be atomic (no double-counting on retries)
  • Stripe webhooks must be idempotent (handle duplicate events)
  • Downgrade mid-cycle: prorate and adjust limits immediately
  • Free tier must work without Stripe (no credit card required)
  • Does NOT include admin dashboard (Phase 4+) or custom enterprise contracts

Acceptance Criteria:

  • Rate limiter returns 429 with retry-after header when exceeded
  • Usage counter accurately tracks pages scraped per API key
  • Free tier enforced at 100 pages/month without credit card
  • Pro tier checkout creates Stripe subscription
  • Subscription webhook updates user tier in database
  • Usage dashboard endpoint returns current period stats

F003 — Production Infrastructure

Purpose: Make the service production-ready with proper database, CI/CD, and deployment.

User-Facing Description: Reliable, always-available service with automated deployments and monitoring.

Technical Requirements:

  • PostgreSQL schema with Alembic migrations
  • Docker multi-stage build (slim runtime image)
  • docker-compose for local development (API + PostgreSQL + Redis)
  • GitHub Actions: lint (ruff), type-check (mypy), test (pytest), build, deploy
  • Production deployment on Railway (or Render)
  • Health check endpoint (/health) and readiness probe
  • Structured JSON logging with request correlation IDs
  • Environment-based configuration (development/staging/production)

Data Requirements:

  • Full PostgreSQL schema implementing all models from F001 and F002
  • Migration history maintained with Alembic

Edge Cases & Constraints:

  • Migration must handle existing data (if any beta users exist)
  • Docker image must include Playwright + Chromium (large, ~500MB)
  • Redis must be persistent for job queue state
  • Does NOT include horizontal scaling (single instance initially)

Acceptance Criteria:

  • docker-compose up starts full local development stack
  • alembic upgrade head applies all migrations cleanly
  • GitHub Actions CI passes on every PR
  • Production deployment triggered on merge to main
  • Health check returns 200 with service status
  • Structured logs include request_id for tracing

F004 — Documentation & Landing Page

Purpose: Enable developer discovery, evaluation, and adoption.

User-Facing Description: Comprehensive API documentation with examples, a landing page explaining what Parsify does, and integration guides.

Technical Requirements:

  • MkDocs documentation site with Material theme
  • Auto-generated API reference from OpenAPI spec
  • Getting started guide with curl, Python, and JavaScript examples
  • Authentication guide
  • Webhook integration guide
  • Landing page with feature overview and pricing
  • SEO-optimized for "documentation scraping API" keywords

Data Requirements:

  • None (static content)

Edge Cases & Constraints:

  • Documentation must stay in sync with API changes
  • Does NOT include blog, changelog (beyond release notes), or community forum

Acceptance Criteria:

  • Documentation site deployed and accessible
  • API reference matches current endpoints
  • Getting started guide works end-to-end (copy-paste and run)
  • Landing page loads in <3s
  • Pricing page shows all tiers with feature comparison

F005 — SDK & Launch

Purpose: Reduce integration friction with native language SDKs.

User-Facing Description: Install the Parsify SDK in your language and start scraping documentation in 3 lines of code.

Technical Requirements:

  • Python SDK auto-generated from OpenAPI (openapi-generator or custom)
  • JavaScript/TypeScript SDK auto-generated from OpenAPI
  • Published to PyPI and npm respectively
  • SDK includes: authentication, job submission, polling, result download
  • Async/await support in both SDKs
  • Type annotations (Python) and TypeScript definitions (JS)

Data Requirements:

  • None (SDK is client-side only)

Edge Cases & Constraints:

  • SDK version must be pinned to API version
  • Breaking API changes require SDK major version bump
  • Does NOT include SDKs for other languages (Go, Ruby, etc.)

Acceptance Criteria:

  • pip install parsify installs Python SDK
  • npm install parsify installs JS SDK
  • SDK examples in documentation work out of the box
  • Both SDKs have >90% API coverage

User Stories & Journeys

Story 1 — Developer: Scrape Documentation for RAG Pipeline

As a backend developer building a RAG pipeline, I want to submit a documentation URL and get back chunked, classified markdown, So that I can feed it directly into my vector database without manual processing.

Journey:

  1. Developer signs up and gets a free-tier API key
  2. Submits POST /api/v1/scrape with {"url": "https://docs.example.com", "max_depth": 3}
  3. Receives job ID and polls status every 30 seconds
  4. Job completes — downloads zip with cleaned markdown chunks + metadata JSON
  5. Feeds chunks into Pinecone/Weaviate with the provided metadata

Happy Path: Job completes in 5-10 minutes, produces well-structured chunks Error Path: Target site blocks scraper → returns clear error with suggestion to try different user-agent or contact support Edge Case: Documentation has 10,000+ pages → job respects max_pages limit, returns partial results with continuation token


Story 2 — DevOps Engineer: Automated Documentation Freshness

As a DevOps engineer, I want to schedule weekly re-scrapes of our vendor documentation, So that our knowledge base stays current without manual intervention.

Journey:

  1. Sets up a cron job that calls POST /api/v1/scrape weekly
  2. Registers a webhook URL for job completion
  3. Webhook fires → triggers vector DB re-indexing pipeline
  4. Monitors usage dashboard to ensure within tier limits

Happy Path: Weekly scrape detects changes, only re-processes modified pages Error Path: Webhook delivery fails → implements retry with exponential backoff Edge Case: Vendor site restructures URLs → scraper handles redirects, logs new URL mappings


Story 3 — Startup CTO: Evaluate Parsify for Team Use

As a startup CTO evaluating documentation processing tools, I want to test Parsify's quality on our documentation sources, So that I can decide whether to adopt it for our AI product.

Journey:

  1. Visits landing page, reads feature overview
  2. Signs up for free tier (no credit card)
  3. Runs 3-4 test scrapes on different documentation sites
  4. Evaluates output quality (chunking, classification, metadata)
  5. If satisfied, upgrades to pro tier via Stripe checkout
  6. Distributes API keys to team members

Happy Path: Free tier's 100 pages is enough for evaluation Error Path: Output quality is poor for specific site → provides feedback via support Edge Case: Team needs 50+ API keys → enterprise tier discussion


Story 4 — Data Engineer: Bulk Process Uploaded Documents

As a data engineer with existing markdown documentation, I want to upload files for cleaning and chunking without scraping, So that I can standardize my existing documentation for AI consumption.

Journey:

  1. Uploads a zip of markdown files to POST /api/v1/process
  2. Specifies processing options (chunk size, classification model, cleaning patterns)
  3. Receives processed output immediately (synchronous for small batches) or via job (async for large batches)
  4. Downloads results with metadata and vector-ready chunks

Happy Path: 50 files processed in under 2 minutes Error Path: Uploaded file has encoding issues → returns error with specific file and line Edge Case: Upload exceeds 100MB → returns 413 with guidance to split


Design System

Design Principles

Since Parsify is primarily an API service, the "design system" applies to:

  1. API response format consistency

  2. Documentation site styling

  3. Landing page design

  4. Developer-First — Every interface decision optimizes for developer experience (clear errors, consistent schemas, copy-paste examples)

  5. Predictable — Same patterns everywhere: consistent error formats, pagination, filtering

  6. Transparent — Usage, costs, and job status are always visible and accurate

API Response Format

{
  "data": { ... },
  "meta": {
    "request_id": "req_abc123",
    "timestamp": "2026-03-22T10:30:00Z"
  }
}

Error format:

{
  "error": {
    "code": "RATE_LIMIT_EXCEEDED",
    "message": "Rate limit exceeded. Try again in 30 seconds.",
    "details": {
      "limit": 60,
      "remaining": 0,
      "reset_at": "2026-03-22T10:31:00Z"
    }
  },
  "meta": {
    "request_id": "req_abc123"
  }
}

Color Palette (Landing Page / Docs)

Role Color Hex Usage
Primary Deep Blue #1e40af CTAs, links, active states
Primary Light Sky Blue #3b82f6 Hover states, highlights
Background White #ffffff Page background
Surface Light Gray #f8fafc Code blocks, cards
Text Primary Slate 900 #0f172a Body text
Text Secondary Slate 500 #64748b Muted text, labels
Success Green #16a34a Success states, "completed"
Warning Amber #d97706 Warnings, "running"
Error Red #dc2626 Errors, "failed"
Code Mono — JetBrains Mono for code blocks

Typography

Element Font Weight Size
H1 Inter 700 36px
H2 Inter 600 28px
H3 Inter 600 22px
Body Inter 400 16px
Code JetBrains Mono 400 14px
API Endpoint JetBrains Mono 600 16px

Brand Voice & Tone

Brand Personality

Technical, direct, and reliable. Like Stripe's documentation: precise without being sterile, helpful without being condescending.

Tone Guidelines

Context Tone Example
Success Brief, factual "Job completed. 247 pages processed in 8m 32s."
Error Helpful, actionable "The target site returned 403 Forbidden. Try setting a custom user-agent header."
Empty State Guiding "No jobs yet. Submit your first scrape request to get started."
Documentation Clear, example-heavy Show the curl command first, explain second.

Writing Guidelines

  • Do: Lead with code examples
  • Do: Use precise technical language
  • Do: State limits explicitly (not "generous free tier" — say "100 pages/month")
  • Don't: Use marketing superlatives ("blazing fast", "revolutionary")
  • Don't: Hide error details behind generic messages

Target Audience Profiles

Primary — Backend Developer (RAG Pipeline Builder)

  • Demographics: 3-10 years experience, works at startup or mid-size company
  • Goals: Build and maintain AI-powered search/documentation systems
  • Pain Points: Manual scraping scripts break, cleaning is tedious, no good managed service
  • Communication Preference: Show the API, show the output, let them evaluate

Secondary — DevOps / Platform Engineer

  • Demographics: 5+ years experience, manages infrastructure
  • Goals: Automate documentation freshness, integrate into CI/CD
  • Pain Points: Custom scraping infrastructure is fragile, monitoring is manual
  • Communication Preference: Docker compose, helm charts, environment variables

Success Metrics

Metric Target How to Measure
Developer Sign-ups 100 in first month User registration count
Free → Pro Conversion 10% within 30 days Stripe subscription events
API Uptime 99.5% Health check monitoring
Job Success Rate >95% Jobs completed / total jobs
Median Job Duration <10 min for 100 pages Job completion timestamps
API Response Time (p95) <200ms (non-scrape) Request logging

Compliance & Legal

Requirement Applicability Implementation
robots.txt Yes Respect target site robots.txt by default
Rate Limiting Yes Self-imposed limits to avoid hammering target sites
Data Retention Yes Job results deleted after 7 days (configurable)
GDPR Partial No PII in scraped content; user accounts use email only
Terms of Service Required Users responsible for legal right to scrape target content
Privacy Policy Required Minimal data collection: email, usage metrics, payment info (Stripe)