Skip to content

Commit e65506e

Browse files
fix: apply DeepSeek V4 pricing (#95)
1 parent 25cff8b commit e65506e

9 files changed

Lines changed: 177 additions & 45 deletions

File tree

‎README.ja.md‎

Lines changed: 3 additions & 3 deletions
Original file line numberDiff line numberDiff line change
@@ -79,7 +79,7 @@ OpenKyrozen はターミナルで動作する**自己学習型 AI エージェ
7979

8080
| プロバイダー | キーの取得 | コスト |
8181
|-------------|-----------|------|
82-
| **DeepSeek** | [platform.deepseek.com](https://platform.deepseek.com) | ~$0.27/100万入力トークン |
82+
| **DeepSeek** | [platform.deepseek.com](https://platform.deepseek.com) | [V4 のピーク/オフピーク・キャッシュ対応料金](https://api-docs.deepseek.com/quick_start/pricing/) |
8383
| **OpenAI** | [platform.openai.com](https://platform.openai.com) | ~$2.50/100万入力トークン |
8484
| **Anthropic (Claude)** | [console.anthropic.com](https://console.anthropic.com) | ~$3.00/100万入力トークン |
8585
| **Google (Gemini)** | [aistudio.google.com](https://aistudio.google.com) | ~$0.15/100万入力トークン |
@@ -250,8 +250,8 @@ Kyrozen はすべてのリクエストを自動分類し、動作を適応させ
250250
エージェントは簡単なタスクと複雑なタスクで異なるモデルを選択します。これらは上書き可能です:
251251

252252
```bash
253-
export KYROZEN_MODEL_SIMPLE=deepseek-chat
254-
export KYROZEN_MODEL_COMPLEX=deepseek-reasoner
253+
export KYROZEN_MODEL_SIMPLE=deepseek-v4-flash
254+
export KYROZEN_MODEL_COMPLEX=deepseek-v4-pro
255255
```
256256

257257
| プロバイダー | 簡単なタスク(デフォルト) | 複雑なタスク(デフォルト) |

‎README.ko.md‎

Lines changed: 3 additions & 3 deletions
Original file line numberDiff line numberDiff line change
@@ -79,7 +79,7 @@ OpenKyrozen은 터미널에서 실행되는 **자기 학습형 AI 에이전트**
7979

8080
| 제공자 | 키 발급 | 비용 |
8181
|--------|--------|------|
82-
| **DeepSeek** | [platform.deepseek.com](https://platform.deepseek.com) | ~$0.27/100만 입력 토큰 |
82+
| **DeepSeek** | [platform.deepseek.com](https://platform.deepseek.com) | [V4 피크/비피크, 캐시 인식 가격](https://api-docs.deepseek.com/quick_start/pricing/) |
8383
| **OpenAI** | [platform.openai.com](https://platform.openai.com) | ~$2.50/100만 입력 토큰 |
8484
| **Anthropic (Claude)** | [console.anthropic.com](https://console.anthropic.com) | ~$3.00/100만 입력 토큰 |
8585
| **Google (Gemini)** | [aistudio.google.com](https://aistudio.google.com) | ~$0.15/100만 입력 토큰 |
@@ -250,8 +250,8 @@ Kyrozen은 모든 요청을 자동 분류하고 동작을 조정합니다:
250250
에이전트는 간단한 작업과 복잡한 작업에 서로 다른 모델을 선택합니다. 재정의할 수 있습니다:
251251

252252
```bash
253-
export KYROZEN_MODEL_SIMPLE=deepseek-chat
254-
export KYROZEN_MODEL_COMPLEX=deepseek-reasoner
253+
export KYROZEN_MODEL_SIMPLE=deepseek-v4-flash
254+
export KYROZEN_MODEL_COMPLEX=deepseek-v4-pro
255255
```
256256

257257
| 제공자 | 간단한 작업 (기본값) | 복잡한 작업 (기본값) |

‎README.md‎

Lines changed: 4 additions & 4 deletions
Original file line numberDiff line numberDiff line change
@@ -80,7 +80,7 @@ Think of it as an AI teammate that gets smarter every time you use it.
8080

8181
| Provider | Get a key | Cost |
8282
|----------|-----------|------|
83-
| **DeepSeek** | [platform.deepseek.com](https://platform.deepseek.com) | ~$0.27/M input tokens |
83+
| **DeepSeek** | [platform.deepseek.com](https://platform.deepseek.com) | [V4 peak/off-peak, cache-aware pricing](https://api-docs.deepseek.com/quick_start/pricing/) |
8484
| **OpenAI** | [platform.openai.com](https://platform.openai.com) | ~$2.50/M input tokens |
8585
| **Anthropic (Claude)** | [console.anthropic.com](https://console.anthropic.com) | ~$3.00/M input tokens |
8686
| **Google (Gemini)** | [aistudio.google.com](https://aistudio.google.com) | ~$0.15/M input tokens |
@@ -260,13 +260,13 @@ Kyrozen automatically classifies every request and adapts its behavior:
260260
The agent picks different models for simple vs complex tasks. You can override these:
261261

262262
```bash
263-
export KYROZEN_MODEL_SIMPLE=deepseek-chat
264-
export KYROZEN_MODEL_COMPLEX=deepseek-reasoner
263+
export KYROZEN_MODEL_SIMPLE=deepseek-v4-flash
264+
export KYROZEN_MODEL_COMPLEX=deepseek-v4-pro
265265
```
266266

267267
| Provider | Simple tasks (default) | Complex tasks (default) |
268268
|----------|----------------------|------------------------|
269-
| DeepSeek | `deepseek-chat` | `deepseek-reasoner` |
269+
| DeepSeek | `deepseek-v4-flash` | `deepseek-v4-pro` |
270270
| OpenAI | `gpt-4o` | `gpt-4o` |
271271
| Anthropic | `claude-sonnet-4-20250514` | `claude-sonnet-4-20250514` |
272272
| Google | `gemini-2.5-flash` | `gemini-2.5-pro` |

‎README.zh-CN.md‎

Lines changed: 3 additions & 3 deletions
Original file line numberDiff line numberDiff line change
@@ -79,7 +79,7 @@ OpenKyrozen 是一款在终端中运行的**自学习 AI 智能体**。与普通
7979

8080
| 服务商 | 获取密钥 | 费用 |
8181
|--------|---------|------|
82-
| **DeepSeek** | [platform.deepseek.com](https://platform.deepseek.com) | ~$0.27/百万输入 token |
82+
| **DeepSeek** | [platform.deepseek.com](https://platform.deepseek.com) | [V4 高峰/低峰、缓存感知计费](https://api-docs.deepseek.com/quick_start/pricing/) |
8383
| **OpenAI** | [platform.openai.com](https://platform.openai.com) | ~$2.50/百万输入 token |
8484
| **Anthropic (Claude)** | [console.anthropic.com](https://console.anthropic.com) | ~$3.00/百万输入 token |
8585
| **Google (Gemini)** | [aistudio.google.com](https://aistudio.google.com) | ~$0.15/百万输入 token |
@@ -246,8 +246,8 @@ Kyrozen 自动对每个请求进行分类并调整行为:
246246
智能体针对简单任务和复杂任务选择不同模型。你可以覆盖这些默认值:
247247

248248
```bash
249-
export KYROZEN_MODEL_SIMPLE=deepseek-chat
250-
export KYROZEN_MODEL_COMPLEX=deepseek-reasoner
249+
export KYROZEN_MODEL_SIMPLE=deepseek-v4-flash
250+
export KYROZEN_MODEL_COMPLEX=deepseek-v4-pro
251251
```
252252

253253
| 服务商 | 简单任务(默认) | 复杂任务(默认) |

‎docs/self-evolution.md‎

Lines changed: 1 addition & 1 deletion
Original file line numberDiff line numberDiff line change
@@ -9,7 +9,7 @@ Historical verification snapshot: `51be33361422e55e1f2f00c33a0e0f8c56132a91`
99
(the post-#54 `main` revision, captured before this #55 documentation-only
1010
update). Snapshot date: 2026-09-04.
1111

12-
Current repository test count at this snapshot: **155 unittest cases**.
12+
Current repository test count at this snapshot: **157 unittest cases**.
1313

1414
## Verified surface
1515

‎event_store.py‎

Lines changed: 14 additions & 3 deletions
Original file line numberDiff line numberDiff line change
@@ -376,6 +376,8 @@ def usage_totals(
376376
pricing_groups = db.execute(
377377
"SELECT provider, pricing_snapshot, "
378378
"SUM(COALESCE(prompt_tokens, 0)) AS prompt_tokens, "
379+
"SUM(COALESCE(cache_hit_tokens, 0)) AS cache_hit_tokens, "
380+
"SUM(COALESCE(cache_miss_tokens, 0)) AS cache_miss_tokens, "
379381
"SUM(COALESCE(completion_tokens, 0)) AS completion_tokens, "
380382
"SUM(COALESCE(cost_picos, 0)) AS cost_picos "
381383
f"FROM usage_attempts WHERE {' AND '.join(clauses)} "
@@ -399,14 +401,23 @@ def exact_cost(groups: list[sqlite3.Row]) -> int:
399401
for row in groups:
400402
snapshot = self._loads(row["pricing_snapshot"], {})
401403
try:
402-
input_rate = max(0, int(snapshot["input_picos_per_million"]))
403404
output_rate = max(0, int(snapshot["output_picos_per_million"]))
405+
cache_miss_rate = snapshot.get("cache_miss_picos_per_million")
406+
input_rate = max(0, int(
407+
cache_miss_rate if cache_miss_rate is not None else snapshot["input_picos_per_million"],
408+
))
404409
except (KeyError, TypeError, ValueError):
405410
# Compatibility for any manually-created legacy ledger row.
406411
fallback += int(row["cost_picos"] or 0)
407412
continue
408-
numerator += (int(row["prompt_tokens"] or 0) * input_rate
409-
+ int(row["completion_tokens"] or 0) * output_rate)
413+
prompt = int(row["prompt_tokens"] or 0)
414+
if "cache_hit_picos_per_million" in snapshot:
415+
hit = min(prompt, max(0, int(row["cache_hit_tokens"] or 0)))
416+
numerator += (hit * max(0, int(snapshot["cache_hit_picos_per_million"]))
417+
+ (prompt - hit) * input_rate)
418+
else:
419+
numerator += prompt * input_rate
420+
numerator += int(row["completion_tokens"] or 0) * output_rate
410421
return fallback + numerator // 1_000_000
411422

412423
totals = normalise(total)

‎main.py‎

Lines changed: 3 additions & 3 deletions
Original file line numberDiff line numberDiff line change
@@ -338,9 +338,9 @@ def _confirm_tool_action(action: str, args: str = "") -> bool:
338338
llm_provider: LLMProvider | None = None
339339

340340
# Backward-compatible aliases (used throughout the codebase)
341-
DEEPSEEK_MODEL_SIMPLE = "deepseek-chat" # set at init time from provider
342-
DEEPSEEK_MODEL_COMPLEX = "deepseek-reasoner"
343-
MODEL_NAME = "deepseek-chat (V4 auto-select)" # updated at init
341+
DEEPSEEK_MODEL_SIMPLE = "deepseek-v4-flash" # set at init time from provider
342+
DEEPSEEK_MODEL_COMPLEX = "deepseek-v4-pro"
343+
MODEL_NAME = "deepseek-v4-flash" # updated at init
344344
# -------- Self-learning feature flags (toggled via /self-learning) --------
345345
# Keep this list as the public feature contract. The dispatcher below binds
346346
# each name to exactly one bounded executor and both CLI and Web use it.

‎providers.py‎

Lines changed: 84 additions & 9 deletions
Original file line numberDiff line numberDiff line change
@@ -23,6 +23,7 @@
2323
from contextlib import contextmanager
2424
from contextvars import ContextVar
2525
from dataclasses import dataclass
26+
from datetime import datetime, timezone
2627
from decimal import Decimal
2728
from typing import Any, Iterator
2829
from abc import ABC, abstractmethod
@@ -34,7 +35,7 @@
3435
# ---------------------------------------------------------------------------
3536

3637
PROVIDER_DEFAULT_MODELS: dict[str, tuple[str, str]] = {
37-
"deepseek": ("deepseek-chat", "deepseek-reasoner"),
38+
"deepseek": ("deepseek-v4-flash", "deepseek-v4-pro"),
3839
"openai": ("gpt-4o", "gpt-4o"),
3940
"anthropic": ("claude-sonnet-4-20250514", "claude-sonnet-4-20250514"),
4041
"google": ("gemini-2.5-flash", "gemini-2.5-pro"),
@@ -81,6 +82,16 @@
8182

8283
_PICOS_PER_DOLLAR = 10**12
8384
_TOKENS_PER_MILLION = 1_000_000
85+
_DEEPSEEK_V4_EFFECTIVE_AT = datetime(2026, 8, 16, 16, tzinfo=timezone.utc)
86+
_DEEPSEEK_V4_MODELS = {
87+
"deepseek-v4-flash": ("0.014", "0.44", "1.32"),
88+
"deepseek-v4-flash-vision-exp": ("0.014", "0.44", "1.32"),
89+
"deepseek-v4-pro": ("0.044", "1.32", "3.96"),
90+
}
91+
_DEEPSEEK_V4_ALIASES = {
92+
"deepseek-chat": "deepseek-v4-flash",
93+
"deepseek-reasoner": "deepseek-v4-flash",
94+
}
8495

8596

8697
@dataclass(frozen=True)
@@ -129,6 +140,35 @@ def _legacy_pricing_snapshot(provider: str) -> tuple[dict[str, Any], int, int]:
129140
}, input_picos, output_picos
130141

131142

143+
def _deepseek_v4_pricing_snapshot(model: str, occurred_at: datetime) -> tuple[dict[str, Any] | None, bool]:
144+
"""Freeze the official DeepSeek V4 rate card selected at a UTC timestamp."""
145+
canonical_model = _DEEPSEEK_V4_ALIASES.get(model.strip().lower(), model.strip().lower())
146+
rates = _DEEPSEEK_V4_MODELS.get(canonical_model)
147+
if rates is None:
148+
return None, False
149+
instant = occurred_at.astimezone(timezone.utc)
150+
peak = instant.weekday() < 5 and (1 <= instant.hour < 4 or 6 <= instant.hour < 10)
151+
multiplier = 1 if peak else 0.5
152+
hit, miss, output = (
153+
int(Decimal(rate) * Decimal(str(multiplier)) * _PICOS_PER_DOLLAR)
154+
for rate in rates
155+
)
156+
current_schedule = instant >= _DEEPSEEK_V4_EFFECTIVE_AT
157+
return {
158+
"version": "deepseek-v4-pricing-2026-08-16",
159+
"source": "https://api-docs.deepseek.com/quick_start/pricing/",
160+
"currency": "USD",
161+
"model": canonical_model,
162+
"priced_at": instant.isoformat(),
163+
"effective_at": _DEEPSEEK_V4_EFFECTIVE_AT.isoformat(),
164+
"billing_window": "peak" if peak else "off_peak",
165+
"cache_hit_picos_per_million": hit,
166+
"cache_miss_picos_per_million": miss,
167+
"output_picos_per_million": output,
168+
"pricing_status": "authoritative" if current_schedule else "estimated_current_schedule",
169+
}, current_schedule
170+
171+
132172
def _usage_integer(usage: dict | None, key: str) -> int | None:
133173
if usage is None or usage.get(key) is None:
134174
return None
@@ -139,30 +179,61 @@ def _usage_integer(usage: dict | None, key: str) -> int | None:
139179

140180

141181
def _track_cost(provider: str, usage: dict | None, *, model: str = "unknown",
142-
latency_ms: int | None = None, completion_state: str = "completed") -> str:
182+
latency_ms: int | None = None, completion_state: str = "completed",
183+
occurred_at: datetime | None = None) -> str:
143184
"""Write one immutable provider attempt; summaries always read this ledger."""
144185
scope = _current_usage_scope()
145186
prompt_tokens = _usage_integer(usage, "prompt_tokens")
146187
completion_tokens = _usage_integer(usage, "completion_tokens")
147188
cache_hit_tokens = _usage_integer(usage, "cache_hit_tokens")
148189
cache_miss_tokens = _usage_integer(usage, "cache_miss_tokens")
190+
if cache_hit_tokens is None:
191+
cache_hit_tokens = _usage_integer(usage, "prompt_cache_hit_tokens")
192+
if cache_miss_tokens is None:
193+
cache_miss_tokens = _usage_integer(usage, "prompt_cache_miss_tokens")
149194
reasoning_tokens = _usage_integer(usage, "reasoning_tokens")
150-
pricing_snapshot, input_rate, output_rate = _legacy_pricing_snapshot(provider)
195+
pricing_snapshot = None
196+
current_schedule = True
197+
if provider == "deepseek":
198+
pricing_snapshot, current_schedule = _deepseek_v4_pricing_snapshot(
199+
model, occurred_at or datetime.now(timezone.utc),
200+
)
201+
if pricing_snapshot is None:
202+
pricing_snapshot, input_rate, output_rate = _legacy_pricing_snapshot(provider)
203+
else:
204+
input_rate = int(pricing_snapshot["cache_miss_picos_per_million"])
205+
output_rate = int(pricing_snapshot["output_picos_per_million"])
151206
cost_picos = None
152207
if prompt_tokens is not None or completion_tokens is not None:
153-
cost_picos = (
154-
((prompt_tokens or 0) * input_rate + (completion_tokens or 0) * output_rate)
155-
// _TOKENS_PER_MILLION
156-
)
208+
prompt = prompt_tokens or 0
209+
if "cache_hit_picos_per_million" in pricing_snapshot:
210+
hit = min(prompt, cache_hit_tokens or 0)
211+
input_cost = (hit * int(pricing_snapshot["cache_hit_picos_per_million"])
212+
+ (prompt - hit) * input_rate)
213+
else:
214+
input_cost = prompt * input_rate
215+
cost_picos = (input_cost + (completion_tokens or 0) * output_rate) // _TOKENS_PER_MILLION
216+
if cache_hit_tokens is None and cache_miss_tokens is None:
217+
cache_status = "unknown"
218+
elif (cache_hit_tokens or 0) and (cache_miss_tokens or 0):
219+
cache_status = "mixed"
220+
elif cache_hit_tokens:
221+
cache_status = "hit"
222+
else:
223+
cache_status = "miss"
224+
usage_status = "unknown" if usage is None else (
225+
"authoritative" if current_schedule and (provider != "deepseek" or cache_status != "unknown")
226+
else "estimated"
227+
)
157228
attempt_id = f"usage_{uuid.uuid4().hex}"
158229
scope.store.record_usage_attempt(
159230
attempt_id=attempt_id, provider=provider, model=model, surface=scope.surface,
160231
user_id=scope.user_id, workspace_id=scope.workspace_id, session_id=scope.session_id,
161-
run_id=scope.run_id, cache_status="unknown", prompt_tokens=prompt_tokens,
232+
run_id=scope.run_id, cache_status=cache_status, prompt_tokens=prompt_tokens,
162233
cache_hit_tokens=cache_hit_tokens, cache_miss_tokens=cache_miss_tokens,
163234
completion_tokens=completion_tokens, reasoning_tokens=reasoning_tokens,
164235
latency_ms=latency_ms, completion_state=completion_state,
165-
usage_status="authoritative" if usage is not None else "unknown",
236+
usage_status=usage_status,
166237
cost_picos=cost_picos, pricing_snapshot=pricing_snapshot,
167238
)
168239
return attempt_id
@@ -349,9 +420,13 @@ def _call():
349420
usage = getattr(response, "usage", None)
350421
usage_dict = None
351422
if usage is not None:
423+
completion_details = getattr(usage, "completion_tokens_details", None)
352424
usage_dict = {
353425
"prompt_tokens": usage.prompt_tokens or 0,
354426
"completion_tokens": usage.completion_tokens or 0,
427+
"prompt_cache_hit_tokens": getattr(usage, "prompt_cache_hit_tokens", None),
428+
"prompt_cache_miss_tokens": getattr(usage, "prompt_cache_miss_tokens", None),
429+
"reasoning_tokens": getattr(completion_details, "reasoning_tokens", None),
355430
}
356431
_track_cost(self.config.provider, usage_dict, model=model,
357432
latency_ms=round((time.monotonic() - started) * 1000))

0 commit comments

Comments
 (0)