# Text and streaming OpenAI-shaped chat completions with durable replay. ## POST /api/generate/text/v1/chat/completions Requires **generate** scope and `Idempotency-Key`. Discover a reviewed text model before submitting. This is a subset of the OpenAI chat-completions shape; it is not a Responses, tools, or multimodal chat endpoint. | Field | Type | Limits / default | | --- | --- | --- | | `model` | string, required | Enabled reviewed text model ID, 1–150 characters. | | `messages` | array, required | 1–60 objects with `role` and string `content`. | | `messages[].role` | string | `system`, `user`, or `assistant`. | | `messages[].content` | string | Up to 30,000 characters each. | | `stream` | boolean | false. | | `max_tokens` | integer | 1–8,192; default 4,096. | | `temperature` | number | Optional, 0–2. | | `top_p` | number | Optional, 0–1. | | `response_format` | object | Optional `{type: "text"}` or `{type: "json_object"}`. | | `maxGems` | integer | Optional conservative admission ceiling, 0–1,000,000. | Request bytes are bounded to 160,000; serialized conversation bytes to 128,000. Unknown fields, tools, provider-routing overrides, remote references, and array-valued message content are rejected. ## Non-streaming example ```bash curl https://letsgen.app/api/generate/text/v1/chat/completions \ -H "Authorization: Bearer $LETSGEN_API_KEY" \ -H 'Content-Type: application/json' \ -H 'Idempotency-Key: my-project-text-001' \ --data '{ "model": "TEXT_MODEL_ID_FROM_DISCOVERY", "messages": [{"role": "user", "content": "Write a short scene about a moonlit harbor."}], "max_tokens": 256, "maxGems": 10 }' ``` HTTP `200` returns `id`, `object: "chat.completion"`, Unix `created`, `model`, `choices` with assistant content and `finish_reason`, and usage token counts. ## Streaming Set `stream: true` and consume `Content-Type: text/event-stream`. Use curl `-N` to disable buffering. SSE `data:` frames contain `chat.completion.chunk` objects. Accumulate `choices[].delta.content`; the final successful chunk includes usage, followed by `data: [DONE]`. A stream can also contain `{error: {code: "REQUEST_UNCERTAIN", message}}`. Treat an error or a stream ending without `[DONE]` as incomplete; never assume HTTP 200 alone proves success. Disconnecting does not cancel provider spending. ## Budgets and replay The API reserves a conservative whole-Gem maximum from the reviewed rates, input size, and requested output limit before inference. Measured usage settles once through the account LLM Gem allowance and consent-gated paid Gems. The charge cannot exceed the admitted reservation. Repeat the exact body, including `stream`, with the same originating key and identity to retrieve a saved response. A completed streaming replay may return the whole saved answer in one chunk instead of reproducing the original chunk timings. Pending or ambiguous responses return `409 REQUEST_UNCERTAIN`; they never dispatch again. See [Budgets and retries](/docs/budgets-and-retries).