Skip to content

Commit 20cabcb

Browse files
committed
feat: add llama-server style generation logging + API response format docs
- Non-streaming: logs prompt=Xt, gen=Yt, speed=Z.ZZt/s | <first 120 chars of output> - Streaming: same format with [stream] tag after speed - Added docs/api-response-formats.md for Aegis-AI integration reference
1 parent e96ec6e commit 20cabcb

2 files changed

Lines changed: 207 additions & 0 deletions

File tree

‎Sources/mlx-server/Server.swift‎

Lines changed: 11 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -990,6 +990,12 @@ func handleChatStreaming(
990990
}
991991
cont.yield("data: [DONE]\n\n")
992992
cont.finish()
993+
// llama-server style generation log
994+
let dur = Date().timeIntervalSince(genStart)
995+
let tokPerSec = dur > 0 ? Double(completionTokenCount) / dur : 0
996+
let preview = String(fullText.prefix(120)).replacingOccurrences(of: "\n", with: " ")
997+
let suffix = fullText.count > 120 ? "..." : ""
998+
print("[mlx-server] prompt=\(promptTokenCount)t, gen=\(completionTokenCount)t, speed=\(String(format: "%.2f", tokPerSec))t/s [stream] | \(preview)\(suffix)")
993999
}
9941000
}
9951001
}
@@ -1047,6 +1053,11 @@ func handleChatNonStreaming(
10471053
await stats.requestFinished(tokens: completionTokenCount, duration: duration)
10481054
await semaphore.signal()
10491055

1056+
// ── llama-server style generation log ──
1057+
let tokPerSec = duration > 0 ? Double(completionTokenCount) / duration : 0
1058+
let outputPreview = fullText.prefix(120).replacingOccurrences(of: "\n", with: " ")
1059+
print("[mlx-server] prompt=\(promptTokenCount)t, gen=\(completionTokenCount)t, speed=\(String(format: "%.2f", tokPerSec))t/s | \(outputPreview)\(fullText.count > 120 ? "..." : "")")
1060+
10501061
// ── Apply stop sequences to final text ──
10511062
var finishReason: String
10521063
switch generationStopReason {

‎docs/api-response-formats.md‎

Lines changed: 196 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,196 @@
1+
# mlx-server API Response Formats
2+
3+
OpenAI-compatible API running at `http://127.0.0.1:5413`.
4+
No authentication required when started without `--api-key`.
5+
6+
---
7+
8+
## 1. Non-Streaming Chat Completion
9+
10+
**Request:**
11+
```bash
12+
curl -X POST http://127.0.0.1:5413/v1/chat/completions \
13+
-H "Content-Type: application/json" \
14+
-d '{
15+
"model": "mlx-community/Qwen3.5-122B-A10B-4bit",
16+
"messages": [{"role": "user", "content": "Hello"}],
17+
"max_tokens": 8192,
18+
"stream": false
19+
}'
20+
```
21+
22+
**Response:**
23+
```json
24+
{
25+
"id": "chatcmpl-<uuid>",
26+
"object": "chat.completion",
27+
"created": 1711746000,
28+
"model": "mlx-community/Qwen3.5-122B-A10B-4bit",
29+
"choices": [{
30+
"index": 0,
31+
"message": {
32+
"role": "assistant",
33+
"content": "The full generated text here..."
34+
},
35+
"finish_reason": "stop"
36+
}],
37+
"usage": {
38+
"prompt_tokens": 22,
39+
"completion_tokens": 2048,
40+
"total_tokens": 2070
41+
}
42+
}
43+
```
44+
45+
**Capture:**
46+
- Text: `.choices[0].message.content`
47+
- Done signal: `.choices[0].finish_reason` → `"stop"` | `"length"` | `"tool_calls"`
48+
49+
---
50+
51+
## 2. Streaming Chat Completion (SSE)
52+
53+
**Request:**
54+
```bash
55+
curl -X POST http://127.0.0.1:5413/v1/chat/completions \
56+
-H "Content-Type: application/json" \
57+
-d '{
58+
"model": "mlx-community/Qwen3.5-122B-A10B-4bit",
59+
"messages": [{"role": "user", "content": "Hello"}],
60+
"max_tokens": 8192,
61+
"stream": true
62+
}'
63+
```
64+
65+
**Response (one line per token):**
66+
```
67+
data: {"id":"chatcmpl-<uuid>","object":"chat.completion.chunk","created":1711746000,"model":"...","choices":[{"index":0,"delta":{"role":"assistant","content":"The "},"finish_reason":null}]}
68+
69+
data: {"id":"chatcmpl-<uuid>","object":"chat.completion.chunk","choices":[{"index":0,"delta":{"content":"full "},"finish_reason":null}]}
70+
71+
data: {"id":"chatcmpl-<uuid>","object":"chat.completion.chunk","choices":[{"index":0,"delta":{},"finish_reason":"stop"}]}
72+
73+
data: [DONE]
74+
```
75+
76+
**Capture pattern:**
77+
1. Each line starts with `data: `
78+
2. Strip `data: ` prefix → parse JSON
79+
3. Accumulate `.choices[0].delta.content` (may be empty string on final chunk)
80+
4. Stop when `.choices[0].finish_reason != null` OR line === `"data: [DONE]"`
81+
82+
---
83+
84+
## 3. Health Check
85+
86+
**Request:**
87+
```bash
88+
curl http://127.0.0.1:5413/health
89+
```
90+
91+
**Response:**
92+
```json
93+
{
94+
"status": "ok",
95+
"model": "mlx-community/Qwen3.5-122B-A10B-4bit",
96+
"memory": {
97+
"total_system_mb": 65536,
98+
"active_mb": 3300,
99+
"peak_mb": 4400,
100+
"cache_mb": 2085
101+
},
102+
"stats": {
103+
"requests_total": 1,
104+
"requests_active": 0,
105+
"avg_tokens_per_sec": 3.81,
106+
"tokens_generated": 2048
107+
}
108+
}
109+
```
110+
111+
**Ready check:** `.status === "ok"` and `.stats.requests_active === 0`
112+
113+
---
114+
115+
## 4. Prometheus Metrics
116+
117+
**Request:**
118+
```bash
119+
curl http://127.0.0.1:5413/metrics
120+
```
121+
122+
**Response (plain text Prometheus format):**
123+
```
124+
mlx_server_requests_total 1
125+
mlx_server_requests_active 0
126+
mlx_server_tokens_generated_total 2048
127+
mlx_server_tokens_per_second 3.81
128+
mlx_server_memory_active_bytes 3462052656
129+
mlx_server_memory_peak_bytes 4577196678
130+
mlx_server_memory_cache_bytes 2185709558
131+
mlx_server_uptime_seconds 652
132+
```
133+
134+
**Generation in progress:** `mlx_server_requests_active > 0` AND `mlx_server_tokens_generated_total == 0` → still prefilling
135+
136+
---
137+
138+
## 5. Aegis-AI Integration (OpenAI SDK)
139+
140+
```javascript
141+
import OpenAI from "openai";
142+
143+
const client = new OpenAI({
144+
baseURL: "http://127.0.0.1:5413/v1",
145+
apiKey: "none", // auth disabled on local server
146+
});
147+
148+
// Non-streaming
149+
const response = await client.chat.completions.create({
150+
model: "mlx-community/Qwen3.5-122B-A10B-4bit",
151+
messages: [{ role: "user", content: "Hello" }],
152+
max_tokens: 8192,
153+
});
154+
const text = response.choices[0].message.content;
155+
const finishReason = response.choices[0].finish_reason; // "stop" | "length"
156+
157+
// Streaming
158+
const stream = await client.chat.completions.create({
159+
model: "mlx-community/Qwen3.5-122B-A10B-4bit",
160+
messages: [{ role: "user", content: "Hello" }],
161+
max_tokens: 8192,
162+
stream: true,
163+
});
164+
let fullText = "";
165+
for await (const chunk of stream) {
166+
fullText += chunk.choices[0]?.delta?.content ?? "";
167+
if (chunk.choices[0]?.finish_reason) break;
168+
}
169+
```
170+
171+
---
172+
173+
## 6. Available Parameters (per-request overrides)
174+
175+
| Parameter | Type | Default | Description |
176+
|---|---|---|---|
177+
| `model` | string | required | Model ID or local path |
178+
| `messages` | array | required | Chat history |
179+
| `max_tokens` | int | 2048 | Max tokens to generate |
180+
| `stream` | bool | false | Enable SSE streaming |
181+
| `temperature` | float | 0.6 | Sampling temperature (0=greedy) |
182+
| `top_p` | float | 1.0 | Nucleus sampling threshold |
183+
| `repetition_penalty` | float | disabled | Repetition penalty factor |
184+
| `stop` | string[] | [] | Stop sequences |
185+
| `response_format` | object | none | `{"type": "json_object"}` for JSON mode |
186+
| `tools` | array | none | Tool/function definitions |
187+
| `enable_thinking` | bool | false | Enable `<think>` reasoning tokens |
188+
189+
---
190+
191+
## 7. Benchmark Results (M5 Pro 64GB, 2026-03-29)
192+
193+
| Model | Strategy | GPU Footprint | tok/s |
194+
|---|---|---|---|
195+
| Qwen3.5-122B-A10B-4bit | SSD Streaming | 4.4 GB peak | **3.81** |
196+
| Qwen2.5-0.5B-Instruct-4bit | Full GPU | 0.3 GB | ~100 |

0 commit comments

Comments
 (0)