Skip to content

feat(cli): caption capture images through any OpenAI-compatible vision endpoint - #4849

Open
FenjuFu wants to merge 1 commit into
heygen-com:mainfrom
FenjuFu:feat/openai-compatible-vision-endpoint
Open

FenjuFu wants to merge 1 commit into
heygen-com:mainfrom
FenjuFu:feat/openai-compatible-vision-endpoint

Conversation

@FenjuFu

@FenjuFu FenjuFu commented Oct 1, 2026

Copy link
Copy Markdown

What

hyperframes capture can describe images through any OpenAI-compatible vision endpoint. Set HYPERFRAMES_VISION_BASE_URL, HYPERFRAMES_VISION_API_KEY and HYPERFRAMES_VISION_MODEL together and that endpoint is used ahead of OpenRouter, Vertex and Gemini.

Why

OpenRouter, Vertex and the Gemini API are all hard to reach from mainland China, and none of them can target a self-hosted model. Those users get catalog-fallback descriptions on every capture even when they have a working vision endpoint.

Related work

Closes #4846

Picks up #1809 by @yuemeng200, closed for merge conflicts after a review that endorsed the approach. This PR is rebuilt on current main, which since gained Vertex (#3561) in the same function, and addresses that review:

  • Half-configured endpoint: if any of the three variables is set, all three are required. Otherwise capture warns with the names of the missing variables (never their values), skips captioning, and reports the vision phase as degraded / internal-error, the same as an unparseable Vertex service account. It never falls back to OpenRouter or Gemini.
  • Helper takes providerName as an argument: openAiCompatibleCaptionOne(providerName, endpoint) is self-contained. OpenRouter now goes through it too, and its URL (https://openrouter.ai/api/v1/chat/completions), Bearer header, request body and error text (OpenRouter request failed with HTTP <status>) are unchanged. The existing OpenRouter tests pass without edits.

How

  • resolveCustomVisionEndpoint() returns unset, incomplete with the missing names, or ready with the endpoint. captionImagesWithGemini and postExtractionPhase.hasVisionCredentials both use it, so the asset-descriptions.md header can't claim vision credentials for a half-set endpoint.
  • Trailing slashes on the base URL are dropped before appending /chat/completions.
  • The no-credentials hint in asset-descriptions.md and the docs (guides/authentication.mdx, packages/cli.mdx) mention the new variables. The docs say to pick a model that accepts image input and to give servers that ignore auth any non-empty key.

Test plan

  • Unit tests added: custom endpoint request (URL, auth header, model, max_tokens, image data URI); priority over OpenRouter and Gemini; failed request counted as a provider failure; key and model without a base URL with GEMINI_API_KEY set (no request, no Gemini client, warning names the variable and does not contain the key, phase degraded); base URL alone with OPENROUTER_API_KEY set; resolveCustomVisionEndpoint unset and incomplete cases.
  • Manual testing performed: called captionImagesWithGemini from this branch with the three variables pointed at iFlytek Astron's Token Plan, a live OpenAI-compatible endpoint (https://maas-token-api.cn-huabei-1.xf-yun.com/v2/, trailing slash included), on two PNGs from docs/images/image-thumbnail-evidence/:
    • xopkimik26 (Kimi-K2.6): 2/2 captioned, e.g. "A dark gray webpage displays white text reading "Timeline JPEG thumbnail resource check" above a horizontal strip of synthetic rainbow-colored JPEG test bars."; outcome all zero
    • xopqwen35397b (Qwen3.5-397B): 2/2 captioned; outcome all zero
    • key and model only, with GEMINI_API_KEY set: no request, warning … must be set together; missing HYPERFRAMES_VISION_BASE_URL., outcome internalError: true
  • Documentation updated
  • Comments follow CONTRIBUTING.md "Comments"

Local checks: vitest run src/capture/contentExtractor.test.ts src/capture/contentExtractor.file-race.test.ts passes 36/36. The 7 new tests fail against main's contentExtractor.ts. oxlint and oxfmt --check are clean on the changed files, and so are check-comment-citations.mjs and comment-ratchet.mjs. tsc over the changed files reports nothing in them. I ran this on Windows without building @hyperframes/core, so the only errors were its generated runtime-inline modules.

…n endpoint

Capture descriptions could only use OpenRouter, Vertex or the Gemini
API, none of which is reliably reachable from mainland China or able
to target a self-hosted model.

Setting HYPERFRAMES_VISION_BASE_URL, HYPERFRAMES_VISION_API_KEY and
HYPERFRAMES_VISION_MODEL together now routes captioning through that
endpoint ahead of the other providers. Setting only some of them warns
with the missing names, skips captioning and reports the phase as
degraded instead of falling back to another provider. OpenRouter goes
through the same request helper with its URL, auth header and error
text unchanged.

Signed-off-by: FenjuFu <fufenjupku@gmail.com>

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Capture descriptions: support any OpenAI-compatible vision endpoint

1 participant