Luma Cloud Docs

For developers

Use your existing SDK.

Python, JavaScript and plain HTTP. Set the Luma Cloud URL and key explicitly for this client.

Complete the quickstart first. These examples read LUMA_CLOUD_API_KEY and LUMA_CLOUD_MODEL from your environment. The model value must be an exact ID returned for your key. Run examples on your backend or computer, never in a public browser bundle.

Python

Install the Python SDK
python -m pip install openai
Streaming chat · Python
import os
from openai import OpenAI

client = OpenAI(
    api_key=os.environ["LUMA_CLOUD_API_KEY"],
    base_url="https://api.lumaos.cloud/v1",
    max_retries=0,
)

stream = client.chat.completions.create(
    model=os.environ["LUMA_CLOUD_MODEL"],
    messages=[{"role": "user", "content": "Reply with hello."}],
    stream=True,
)
for chunk in stream:
    if chunk.choices:
        print(chunk.choices[0].delta.content or "", end="", flush=True)
print()

Save as luma_example.py and run python luma_example.py in the terminal where the environment variables are set. Use a virtual environment for your project.

JavaScript / TypeScript

Install the JavaScript SDK
npm install openai
Streaming chat · Node.js
import OpenAI from "openai";

const apiKey = process.env.LUMA_CLOUD_API_KEY;
const model = process.env.LUMA_CLOUD_MODEL;
if (!apiKey || !model) throw new Error("Set LUMA_CLOUD_API_KEY and LUMA_CLOUD_MODEL");

const client = new OpenAI({
  apiKey,
  baseURL: "https://api.lumaos.cloud/v1",
  maxRetries: 0,
});
const stream = await client.chat.completions.create({
  model,
  messages: [{ role: "user", content: "Reply with hello." }],
  stream: true,
});
for await (const chunk of stream) {
  process.stdout.write(chunk.choices[0]?.delta.content ?? "");
}
process.stdout.write("\n");

For JavaScript, save as luma-example.mjs and run node luma-example.mjs. The same client setup works in a server-side TypeScript project.

Using Responses instead

Reuse the client from the corresponding example, replacing the streaming chat call with:

Responses · Python
response = client.responses.create(
    model=os.environ["LUMA_CLOUD_MODEL"],
    input="Reply with hello.",
)
print(response.output_text)
Responses · JavaScript
const response = await client.responses.create({
  model,
  input: "Reply with hello.",
});
console.log(response.output_text);

These examples return a complete response. To stream Responses, set stream: true and handle response text delta, completion and failure events according to your SDK.

Without an SDK

Streaming HTTP · macOS / Linux
curl --fail-with-body -N "https://api.lumaos.cloud/v1/chat/completions" \
  -H "Authorization: Bearer $LUMA_CLOUD_API_KEY" \
  -H "Content-Type: application/json" \
  --data @- <<JSON
{
  "model": "$LUMA_CLOUD_MODEL",
  "messages": [{"role": "user", "content": "Reply with hello."}],
  "stream": true
}
JSON

For Windows, use the PowerShell 7 environment setup in the quickstart, then:

HTTP · PowerShell
$headers = @{ Authorization = "Bearer $env:LUMA_CLOUD_API_KEY" }
$models = Invoke-RestMethod "$env:LUMA_CLOUD_BASE_URL/models" -Headers $headers
$models.data | Select-Object id

$env:LUMA_CLOUD_MODEL = Read-Host "Model ID from the list above"
$body = @{
  model = $env:LUMA_CLOUD_MODEL
  messages = @(@{ role = "user"; content = "Reply with hello." })
  stream = $false
} | ConvertTo-Json -Depth 5
$result = Invoke-RestMethod "$env:LUMA_CLOUD_BASE_URL/chat/completions" -Method Post -Headers $headers -ContentType "application/json" -Body $body
$result.choices[0].message.content

Using another language or framework

Use its OpenAI-compatible client with an explicit base URL and Bearer key. Select Chat Completions or Responses to match the client. For LangChain, LlamaIndex, Vercel AI SDK or your own HTTP client, verify the model adapter’s protocol before enabling tools.

Do not enable embeddings, FIM autocomplete, image generation or hosted tools simply because the framework exposes them. A chat model is not automatically an embedding or autocomplete model.

Examples disable automatic SDK retries to keep the first check easy to understand. Add bounded retries for your application after reviewing errors and rate limits. Never blindly repeat side-effecting tool actions.

Official SDK references

Examples checked against the official SDK documentation on September 28, 2026.