|
| 1 | +# mlx-server API Response Formats |
| 2 | + |
| 3 | +OpenAI-compatible API running at `http://127.0.0.1:5413`. |
| 4 | +No authentication required when started without `--api-key`. |
| 5 | + |
| 6 | +--- |
| 7 | + |
| 8 | +## 1. Non-Streaming Chat Completion |
| 9 | + |
| 10 | +**Request:** |
| 11 | +```bash |
| 12 | +curl -X POST http://127.0.0.1:5413/v1/chat/completions \ |
| 13 | + -H "Content-Type: application/json" \ |
| 14 | + -d '{ |
| 15 | + "model": "mlx-community/Qwen3.5-122B-A10B-4bit", |
| 16 | + "messages": [{"role": "user", "content": "Hello"}], |
| 17 | + "max_tokens": 8192, |
| 18 | + "stream": false |
| 19 | + }' |
| 20 | +``` |
| 21 | + |
| 22 | +**Response:** |
| 23 | +```json |
| 24 | +{ |
| 25 | + "id": "chatcmpl-<uuid>", |
| 26 | + "object": "chat.completion", |
| 27 | + "created": 1711746000, |
| 28 | + "model": "mlx-community/Qwen3.5-122B-A10B-4bit", |
| 29 | + "choices": [{ |
| 30 | + "index": 0, |
| 31 | + "message": { |
| 32 | + "role": "assistant", |
| 33 | + "content": "The full generated text here..." |
| 34 | + }, |
| 35 | + "finish_reason": "stop" |
| 36 | + }], |
| 37 | + "usage": { |
| 38 | + "prompt_tokens": 22, |
| 39 | + "completion_tokens": 2048, |
| 40 | + "total_tokens": 2070 |
| 41 | + } |
| 42 | +} |
| 43 | +``` |
| 44 | + |
| 45 | +**Capture:** |
| 46 | +- Text: `.choices[0].message.content` |
| 47 | +- Done signal: `.choices[0].finish_reason` → `"stop"` | `"length"` | `"tool_calls"` |
| 48 | + |
| 49 | +--- |
| 50 | + |
| 51 | +## 2. Streaming Chat Completion (SSE) |
| 52 | + |
| 53 | +**Request:** |
| 54 | +```bash |
| 55 | +curl -X POST http://127.0.0.1:5413/v1/chat/completions \ |
| 56 | + -H "Content-Type: application/json" \ |
| 57 | + -d '{ |
| 58 | + "model": "mlx-community/Qwen3.5-122B-A10B-4bit", |
| 59 | + "messages": [{"role": "user", "content": "Hello"}], |
| 60 | + "max_tokens": 8192, |
| 61 | + "stream": true |
| 62 | + }' |
| 63 | +``` |
| 64 | + |
| 65 | +**Response (one line per token):** |
| 66 | +``` |
| 67 | +data: {"id":"chatcmpl-<uuid>","object":"chat.completion.chunk","created":1711746000,"model":"...","choices":[{"index":0,"delta":{"role":"assistant","content":"The "},"finish_reason":null}]} |
| 68 | +
|
| 69 | +data: {"id":"chatcmpl-<uuid>","object":"chat.completion.chunk","choices":[{"index":0,"delta":{"content":"full "},"finish_reason":null}]} |
| 70 | +
|
| 71 | +data: {"id":"chatcmpl-<uuid>","object":"chat.completion.chunk","choices":[{"index":0,"delta":{},"finish_reason":"stop"}]} |
| 72 | +
|
| 73 | +data: [DONE] |
| 74 | +``` |
| 75 | + |
| 76 | +**Capture pattern:** |
| 77 | +1. Each line starts with `data: ` |
| 78 | +2. Strip `data: ` prefix → parse JSON |
| 79 | +3. Accumulate `.choices[0].delta.content` (may be empty string on final chunk) |
| 80 | +4. Stop when `.choices[0].finish_reason != null` OR line === `"data: [DONE]"` |
| 81 | + |
| 82 | +--- |
| 83 | + |
| 84 | +## 3. Health Check |
| 85 | + |
| 86 | +**Request:** |
| 87 | +```bash |
| 88 | +curl http://127.0.0.1:5413/health |
| 89 | +``` |
| 90 | + |
| 91 | +**Response:** |
| 92 | +```json |
| 93 | +{ |
| 94 | + "status": "ok", |
| 95 | + "model": "mlx-community/Qwen3.5-122B-A10B-4bit", |
| 96 | + "memory": { |
| 97 | + "total_system_mb": 65536, |
| 98 | + "active_mb": 3300, |
| 99 | + "peak_mb": 4400, |
| 100 | + "cache_mb": 2085 |
| 101 | + }, |
| 102 | + "stats": { |
| 103 | + "requests_total": 1, |
| 104 | + "requests_active": 0, |
| 105 | + "avg_tokens_per_sec": 3.81, |
| 106 | + "tokens_generated": 2048 |
| 107 | + } |
| 108 | +} |
| 109 | +``` |
| 110 | + |
| 111 | +**Ready check:** `.status === "ok"` and `.stats.requests_active === 0` |
| 112 | + |
| 113 | +--- |
| 114 | + |
| 115 | +## 4. Prometheus Metrics |
| 116 | + |
| 117 | +**Request:** |
| 118 | +```bash |
| 119 | +curl http://127.0.0.1:5413/metrics |
| 120 | +``` |
| 121 | + |
| 122 | +**Response (plain text Prometheus format):** |
| 123 | +``` |
| 124 | +mlx_server_requests_total 1 |
| 125 | +mlx_server_requests_active 0 |
| 126 | +mlx_server_tokens_generated_total 2048 |
| 127 | +mlx_server_tokens_per_second 3.81 |
| 128 | +mlx_server_memory_active_bytes 3462052656 |
| 129 | +mlx_server_memory_peak_bytes 4577196678 |
| 130 | +mlx_server_memory_cache_bytes 2185709558 |
| 131 | +mlx_server_uptime_seconds 652 |
| 132 | +``` |
| 133 | + |
| 134 | +**Generation in progress:** `mlx_server_requests_active > 0` AND `mlx_server_tokens_generated_total == 0` → still prefilling |
| 135 | + |
| 136 | +--- |
| 137 | + |
| 138 | +## 5. Aegis-AI Integration (OpenAI SDK) |
| 139 | + |
| 140 | +```javascript |
| 141 | +import OpenAI from "openai"; |
| 142 | + |
| 143 | +const client = new OpenAI({ |
| 144 | + baseURL: "http://127.0.0.1:5413/v1", |
| 145 | + apiKey: "none", // auth disabled on local server |
| 146 | +}); |
| 147 | + |
| 148 | +// Non-streaming |
| 149 | +const response = await client.chat.completions.create({ |
| 150 | + model: "mlx-community/Qwen3.5-122B-A10B-4bit", |
| 151 | + messages: [{ role: "user", content: "Hello" }], |
| 152 | + max_tokens: 8192, |
| 153 | +}); |
| 154 | +const text = response.choices[0].message.content; |
| 155 | +const finishReason = response.choices[0].finish_reason; // "stop" | "length" |
| 156 | + |
| 157 | +// Streaming |
| 158 | +const stream = await client.chat.completions.create({ |
| 159 | + model: "mlx-community/Qwen3.5-122B-A10B-4bit", |
| 160 | + messages: [{ role: "user", content: "Hello" }], |
| 161 | + max_tokens: 8192, |
| 162 | + stream: true, |
| 163 | +}); |
| 164 | +let fullText = ""; |
| 165 | +for await (const chunk of stream) { |
| 166 | + fullText += chunk.choices[0]?.delta?.content ?? ""; |
| 167 | + if (chunk.choices[0]?.finish_reason) break; |
| 168 | +} |
| 169 | +``` |
| 170 | +
|
| 171 | +--- |
| 172 | +
|
| 173 | +## 6. Available Parameters (per-request overrides) |
| 174 | +
|
| 175 | +| Parameter | Type | Default | Description | |
| 176 | +|---|---|---|---| |
| 177 | +| `model` | string | required | Model ID or local path | |
| 178 | +| `messages` | array | required | Chat history | |
| 179 | +| `max_tokens` | int | 2048 | Max tokens to generate | |
| 180 | +| `stream` | bool | false | Enable SSE streaming | |
| 181 | +| `temperature` | float | 0.6 | Sampling temperature (0=greedy) | |
| 182 | +| `top_p` | float | 1.0 | Nucleus sampling threshold | |
| 183 | +| `repetition_penalty` | float | disabled | Repetition penalty factor | |
| 184 | +| `stop` | string[] | [] | Stop sequences | |
| 185 | +| `response_format` | object | none | `{"type": "json_object"}` for JSON mode | |
| 186 | +| `tools` | array | none | Tool/function definitions | |
| 187 | +| `enable_thinking` | bool | false | Enable `<think>` reasoning tokens | |
| 188 | +
|
| 189 | +--- |
| 190 | +
|
| 191 | +## 7. Benchmark Results (M5 Pro 64GB, 2026-03-29) |
| 192 | +
|
| 193 | +| Model | Strategy | GPU Footprint | tok/s | |
| 194 | +|---|---|---|---| |
| 195 | +| Qwen3.5-122B-A10B-4bit | SSD Streaming | 4.4 GB peak | **3.81** | |
| 196 | +| Qwen2.5-0.5B-Instruct-4bit | Full GPU | 0.3 GB | ~100 | |
0 commit comments