For developers
Use your existing SDK.
Python, JavaScript and plain HTTP. Set the Luma Cloud URL and key explicitly for this client.
Complete the quickstart first. These examples read LUMA_CLOUD_API_KEY and LUMA_CLOUD_MODEL from your environment. The model value must be an exact ID returned for your key. Run examples on your backend or computer, never in a public browser bundle.
Python
python -m pip install openaiimport os
from openai import OpenAI
client = OpenAI(
api_key=os.environ["LUMA_CLOUD_API_KEY"],
base_url="https://api.lumaos.cloud/v1",
max_retries=0,
)
stream = client.chat.completions.create(
model=os.environ["LUMA_CLOUD_MODEL"],
messages=[{"role": "user", "content": "Reply with hello."}],
stream=True,
)
for chunk in stream:
if chunk.choices:
print(chunk.choices[0].delta.content or "", end="", flush=True)
print()Save as luma_example.py and run python luma_example.py in the terminal where the environment variables are set. Use a virtual environment for your project.
JavaScript / TypeScript
npm install openaiimport OpenAI from "openai";
const apiKey = process.env.LUMA_CLOUD_API_KEY;
const model = process.env.LUMA_CLOUD_MODEL;
if (!apiKey || !model) throw new Error("Set LUMA_CLOUD_API_KEY and LUMA_CLOUD_MODEL");
const client = new OpenAI({
apiKey,
baseURL: "https://api.lumaos.cloud/v1",
maxRetries: 0,
});
const stream = await client.chat.completions.create({
model,
messages: [{ role: "user", content: "Reply with hello." }],
stream: true,
});
for await (const chunk of stream) {
process.stdout.write(chunk.choices[0]?.delta.content ?? "");
}
process.stdout.write("\n");For JavaScript, save as luma-example.mjs and run node luma-example.mjs. The same client setup works in a server-side TypeScript project.
Using Responses instead
Reuse the client from the corresponding example, replacing the streaming chat call with:
response = client.responses.create(
model=os.environ["LUMA_CLOUD_MODEL"],
input="Reply with hello.",
)
print(response.output_text)const response = await client.responses.create({
model,
input: "Reply with hello.",
});
console.log(response.output_text);These examples return a complete response. To stream Responses, set stream: true and handle response text delta, completion and failure events according to your SDK.
Without an SDK
curl --fail-with-body -N "https://api.lumaos.cloud/v1/chat/completions" \
-H "Authorization: Bearer $LUMA_CLOUD_API_KEY" \
-H "Content-Type: application/json" \
--data @- <<JSON
{
"model": "$LUMA_CLOUD_MODEL",
"messages": [{"role": "user", "content": "Reply with hello."}],
"stream": true
}
JSONFor Windows, use the PowerShell 7 environment setup in the quickstart, then:
$headers = @{ Authorization = "Bearer $env:LUMA_CLOUD_API_KEY" }
$models = Invoke-RestMethod "$env:LUMA_CLOUD_BASE_URL/models" -Headers $headers
$models.data | Select-Object id
$env:LUMA_CLOUD_MODEL = Read-Host "Model ID from the list above"
$body = @{
model = $env:LUMA_CLOUD_MODEL
messages = @(@{ role = "user"; content = "Reply with hello." })
stream = $false
} | ConvertTo-Json -Depth 5
$result = Invoke-RestMethod "$env:LUMA_CLOUD_BASE_URL/chat/completions" -Method Post -Headers $headers -ContentType "application/json" -Body $body
$result.choices[0].message.contentUsing another language or framework
Use its OpenAI-compatible client with an explicit base URL and Bearer key. Select Chat Completions or Responses to match the client. For LangChain, LlamaIndex, Vercel AI SDK or your own HTTP client, verify the model adapter’s protocol before enabling tools.
Do not enable embeddings, FIM autocomplete, image generation or hosted tools simply because the framework exposes them. A chat model is not automatically an embedding or autocomplete model.
Examples disable automatic SDK retries to keep the first check easy to understand. Add bounded retries for your application after reviewing errors and rate limits. Never blindly repeat side-effecting tool actions.
Official SDK references
Examples checked against the official SDK documentation on September 28, 2026.
