API docs
Use the API at https://api.chigyu.ai.
Endpoints
GET /health — gateway health
Gateway liveness. Should return "ok", otherwise the server is down.
curl https://api.chigyu.ai/health
Response · 200
{"status": "ok"}
GET /endpoints — LLM endpoint availability
LLM endpoint availability. Serving engines are swapped out frequently and sometimes there might be multiple online, this will list them.
curl https://api.chigyu.ai/endpoints \
-H "Authorization: Bearer $CHIGYU_API_KEY"
Response · 200
{
"endpoints": [{
"name": "radiance-1",
"model": "qwen3.8-27b",
"online": true,
"checked_at": 1791169353
}]
}
GET /v1/models — online models
Standard OpenAI-compatible /models endpoint.
curl https://api.chigyu.ai/v1/models \
-H "Authorization: Bearer $CHIGYU_API_KEY"
Response · 200
{
"object": "list",
"data": [{
"id": "qwen3.8-27b",
"object": "model",
"created": 0,
"owned_by": "lutetai"
}]
}
Use the returned model ID in inference requests.
POST /v1/chat/completions — model serving
Standard OpenAI-compatible /chat/completions endpoint.
curl https://api.chigyu.ai/v1/chat/completions \
-H "Authorization: Bearer $CHIGYU_API_KEY" \
-H 'Content-Type: application/json' \
-d '{
"model": "qwen3.8-27b",
"messages": [{"role": "user", "content": "Say hello."}],
"reasoning_effort": "low",
"max_tokens": 1024
}'
Response · 200 (abridged)
{
"object": "chat.completion",
"model": "qwen3.8-27b",
"choices": [{
"index": 0,
"message": {"role": "assistant", "content": "Hello!"},
"finish_reason": "stop"
}],
"usage": {"prompt_tokens": 19, "completion_tokens": 2, "total_tokens": 21}
}
Streaming
curl -N https://api.chigyu.ai/v1/chat/completions \
-H "Authorization: Bearer $CHIGYU_API_KEY" \
-H 'Content-Type: application/json' \
-d '{
"model": "qwen3.8-27b",
"messages": [{"role": "user", "content": "Say hello."}],
"reasoning_effort": "low",
"max_tokens": 1024,
"stream": true
}'
Response · 200 · text/event-stream (abridged)
data: {"choices":[{"index":0,"delta":{"content":"Hello!"}}]}
data: {"choices":[],"usage":{"prompt_tokens":19,"completion_tokens":2,"total_tokens":21}}
data: [DONE]
POST /v1/completions — text prompts
Standard OpenAI-compatible /completions endpoint.
curl https://api.chigyu.ai/v1/completions \
-H "Authorization: Bearer $CHIGYU_API_KEY" \
-H 'Content-Type: application/json' \
-d '{"model":"qwen3.8-27b","prompt":"The capital of France is","max_tokens":32}'
Response · 200 (abridged)
{
"object": "text_completion",
"model": "qwen3.8-27b",
"choices": [{"index": 0, "text": " Paris.", "finish_reason": "stop"}]
}
Raw text completion; use chat/completions for conversations. Streaming is also supported.
Python · OpenAI SDK
pip install openai
import os
from openai import OpenAI
client = OpenAI(
base_url="https://api.chigyu.ai/v1",
api_key=os.environ["CHIGYU_API_KEY"],
)
stream = client.chat.completions.create(
model="qwen3.8-27b",
messages=[{"role": "user", "content": "Say hello."}],
reasoning_effort="low",
max_tokens=1024,
stream=True,
)
for chunk in stream:
if chunk.choices:
print(chunk.choices[0].delta.content or "", end="", flush=True)
print()
OpenCode V2
Add this provider to ~/.config/opencode/opencode.json.
{
"$schema": "https://opencode.ai/config.json",
"providers": {
"chigyu": {
"name": "chigyu.ai (lutetai)",
"env": ["CHIGYU_API_KEY"],
"package": "@opencode/ai/providers/openai-compatible",
"settings": {"baseURL": "https://api.chigyu.ai/v1"},
"models": {
"qwen3.8-27b": {
"name": "Qwen3.8 27B (chigyu.ai)",
"settings": {"reasoningEffort": "xhigh"},
"variants": [
{"id": "low", "settings": {"reasoningEffort": "low"}},
{"id": "medium", "settings": {"reasoningEffort": "medium"}},
{"id": "high", "settings": {"reasoningEffort": "xhigh"}},
{"id": "xhigh", "settings": {"reasoningEffort": "xhigh"}}
],
"capabilities": {
"tools": true,
"input": ["text", "image"],
"output": ["text"]
},
"limit": {"context": 262144, "output": 32768}
}
}
}
}
}
Use /connect to enter your API key, or supply CHIGYU_API_KEY
to the OpenCode server.
Then select chigyu/qwen3.8-27b in /models.
The high preset maps to xhigh because this backend does not accept high.
Serving engines
We use various open source inference engines for model serving.
Currently, we are just using a slightly modified vLLM Radiance.