Skip to content

Add genomic-intelligence skill - #43

Open
boldakov wants to merge 13 commits into
FreedomIntelligence:mainfrom
genomicintelligence:add-genomic-intelligence
Open

boldakov wants to merge 13 commits into
FreedomIntelligence:mainfrom
genomicintelligence:add-genomic-intelligence

Conversation

@boldakov

Copy link
Copy Markdown

Add genomic-intelligence skill

Adds a skill wrapping Genomic Intelligence (GI), a set of hosted transformer DNA
language models that predict regulatory features, gene structure, and expression
directly from DNA sequence. No local GPU or model weights are needed; the skill is
a thin client over a hosted, versioned inference API.

What it does

Give it a gene symbol, a genomic region, or a DNA/FASTA sequence and it returns
structured predictions across six tasks, plus a composite:

  • promoter: promoter regions (sliding window)
  • splice: donor/acceptor splice sites
  • enhancer: developmental and housekeeping enhancer activity
  • chromatin: chromatin state across hundreds of tracks
  • expression: sequence to expression, log(TPM+1)
  • annotation: de-novo gene/transcript annotation
  • composite: find the genes in a region and predict each one's expression

Shape

This mirrors the existing hosted-API skills (alphafold-database,
chembl-database, pubchem-database): a SKILL.md with name and description
frontmatter plus a references/ directory, using inline requests with no
vendored client.

There are two ways to call GI. The REST /v1 API needs a GI_API_KEY bearer. The
hosted MCP server at https://mcp.genomicintelligence.ai/mcp works keyless
against a public demo quota, with the key optional for a higher quota.

references/ documents the tasks, the REST contract and auth, the MCP tool flow,
and Ensembl-based sequence acquisition, including the exact 9,198 bp TSS-centred
window the expression model requires.

Registration

Appended "./skills/genomic-intelligence" to the skills array in
openclaw.plugin.json, in alphabetical position. Added a row to the README skills
table under Scientific Databases (Genomics & Variants) and bumped the Skills
Overview counts from 43 to 44 for that category and from 869 to 870 for the total.

Validation

python scripts/validate_skill.py skills/genomic-intelligence reports "Skill is
valid!" and python scripts/validate_skill.py openclaw.plugin.json reports
"Manifest valid, found 870 skills". No secrets are committed; the key is read from
the GI_API_KEY environment variable.

The keyless MCP path was verified against the live endpoint: an anonymous
Streamable HTTP session with no Authorization header completes initialize and
returns a real predict_promoter result.

Notes

For research and development use, not clinical or diagnostic decisions.

Docs: https://docs.genomicintelligence.ai

boldakov and others added 12 commits July 23, 2026 23:24
Wrap Genomic Intelligence's hosted DNA-sequence models (promoter, splice,
enhancer, chromatin, expression, annotation, plus a find-genes-then-expression
composite) over the REST /v1 API and the keyless hosted MCP server. SKILL.md +
references only, no vendored client — mirroring the alphafold-database skill.
Registered in openclaw.plugin.json and the README skills table.
fetch_ensembl_sequence takes gene (fetch_region handles coordinates),
load_demo_sequence requires name, and find_genes /
find_genes_and_predict_expression take a sequence handle rather than a region.
The annotation task is find_genes; there is no predict_annotation and no
load_local_fasta on the hosted server. Also drop dangling references/errors.md
pointers, add the 422 validation_failed case, and state the 9,198 bp window
consistently.
869 appears four times in README.md; the initial commit bumped only the Skills
Overview total, leaving the badge, tagline and intro contradicting it. Bump all
of them, and mirror everything in README_zh.md (counts plus the
genomic-intelligence row), which the repo keeps in sync with README.md.
…8-500,000 bp)

Genomic Intelligence shipped a new expression contract to production on
2026-08-18 (gpu_service 2026.08.18.1). Expression now has its own
published OpenAPI operation (POST /v1/tasks/expression/predict, schema
ExpressionPredictRequest) instead of the generic templated one, so the
skill's description of it was wrong in four ways:

- the sequence bound is 9,198-500,000 bp, not "exactly 9,198 bp";
- a new required field tss_index (0-based, into the whitespace-stripped
  sequence, 4599 <= tss_index <= len-4599) selects the window the server
  slices, and is mandatory unless the sequence is exactly 9,198 bp;
- options is now a closed object whose only, required, key is
  description;
- violations are 422 validation_failed, never a padded/truncated 200.

Also documents the response's new tss_index / scored_window /
submitted_sequence_length provenance fields, and the gotcha that an
in-range but wrong tss_index returns a confident 200 for the wrong
window. The five other tasks and the composite workflow are unchanged.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
… options)

Correct the skill against the settled /v1 contract (DEV 2026.08.19.4):

- Replace "all six tasks share one shape / one PredictRequest". Six literal
  predict operations are published, one per task, each with its own request
  schema, minLength and closed options object. The URLs are unchanged.
- Replace the "1-500,000 bp" length bound with the real per-task floors:
  promoter 300, splice 100, enhancer 50, chromatin 200, annotation 1,000,
  expression 9,198, composite 1,000; 500,000 cap throughout.
- Add the floor-is-not-regime rule: a request above the floor but below the
  model's bio_spec.context_window_bp is accepted and scored against a padded
  window. Enhancer's 50 bp bound vs its 249 bp context window is the sharp case.
- Correct "413 sequence too long": over-length is 422 validation_failed. 413
  means the 16 MiB raw-body cap or the composite's sync_too_large.
- Document typed, closed per-task options (an unknown key is a hard 422
  extra_forbidden), the closed 21-value error code enum, the defensive-details
  rule, the Prefer header on all six operations, and the published composite
  REST workflow.
- Note that tss_index errors report at loc ["body"], never body.tss_index, so
  clients must match on error.code.
- Document the new bio_spec fields (request_max_bp, context_window_bp,
  trained_window_bp), that max_seq_length_bp is legacy/ambiguous, that there is
  no strand_sensitive flag, and that GET /v1/tasks/{task}/models returns a flat
  object rather than the {data, meta} envelope.
- Add an unknown task -> 404 not_found note and a schema-version caveat, since
  production may still serve 2026.08.18.1 until this build is promoted.

Verified against the live DEV schema (info.version 2026.08.19.4 (a5b1f88),
11 operations, no PredictRequest component).
Expression reports `context_window_bp: null`; 9,198 bp is
`trained_window_bp`. The SKILL.md and references/tasks.md tables listed
9,198 under a "Model context window" column, conflating the two.
…property

Every predict operation declares the Prefer header and both 200 and 202.
Verified live against 2026.08.19.4: promoter submitted with
Prefer: respond-async returns 202 and polls to a normal payload; annotation
with no Prefer header returns 200 synchronously.

Relabels the task table's "Mode" column "Recommended mode" and states the rule
under the table, so the latency guidance stays without asserting a constraint
the API never had.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Alexander Boldakov <alexboldakov@gmail.com>
error.request_id is not always populated — 413 sync_too_large omits it from
the body. The X-Request-Id header is set on every response, so read the
header as a fallback.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Alexander Boldakov <alexboldakov@gmail.com>
…08.19.5

Today's GI ship landed after this text was written, so several claims went
stale within hours:

- 422 details is the declared {"errors": [...]} object, not a bare array.
  Defensive handling of either shape is kept; only the prose was wrong.
- error.request_id is populated on every error path, including both 413
  variants; success envelopes carry meta.request_id.
- bio_spec.max_seq_length_bp is retired and absent from the live response;
  request_max_bp is the cap.
- Version notes hedging that PROD may still serve an older build are moot.
- Sync/async is a per-request delivery choice, never a task property; the
  composite's 50,000 bp sync limit is the only forced case.

Also renames the client's synthesized non_json error code to http_error,
which is in the published 21-value enum; client-origin errors remain
distinguishable by carrying no request_id.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Alexander Boldakov <alexboldakov@gmail.com>
…omments

Audited against AUTHORING.md. Removes every pointer at the private service
codebase (retired `gpu_service/*` paths, internal constants and symbols) and
the internal release stamps that went with them, restating each as what the
Genomic Intelligence API serves, with the published OpenAPI document as the
authority. Also normalizes the research-use wording, removes pinned model IDs
and prediction values from prose, and makes comment voice consistent.

No contract facts changed; every bound, field and status code re-verified
against the live schema.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
splice is strand-specific and the wrong strand fails silently: a
reverse-complemented sequence returns plausible sites at different positions,
often at high confidence, so neither the score nor the count reveals the
orientation. Without this stated, the natural inference from "strand-specific"
is that a bad strand shows up as a weak or empty result, which is backwards.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…unction base

The reference documented data.sites' start/end without saying what the pair
means. It bounds one variable-width tokenizer token (4-10 bp observed), comes
with a token_index, and the exon/intron junction sits inside it. Reading either
coordinate as the boundary base is wrong by up to about 10 bp in either
orientation, and the gff3 output carries the same span in columns 4-5, so a
saved track intersected against reference annotation inherits the error.

Verified against the live contract (info.version 2026.08.20.11) and the pinned
HBB fixture, whose eight sites span 4, 5, 6, 7 and 8 bp.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
@boldakov

Copy link
Copy Markdown
Author

Could a maintainer approve the workflow runs here? As a fork PR they need a manual approval, and all 7 runs since the PR opened are at action_required, so CI has never executed against this branch. Happy to fix whatever it reports once it runs.

data.input.sequence_length was removed from the API at contract revision 13;
the docs described it as the scored 9,198 bp.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LA7HntiXoc7ZpQz8boGk6b
@boldakov

boldakov commented Sep 3, 2026

Copy link
Copy Markdown
Author

Small docs correction: the API dropped data.input.sequence_length (it was the 9,198 bp scored window and got mistaken for the submitted length). The submitted length is meta.sequence_length, and the two references now say so. No code change.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant