# AnythingLLM + Luma Cloud Use Luma Cloud as the language model for your workspace. API connection: Chat Completions. Instructions checked: 2026-09-28. Client capabilities and versions may vary; follow the verification steps below. This is an API setup guide. SuperGPT desktop installation and sign-in are separate; use a customer API key for the client described here. ## Before you start - Create a named key in Dashboard → API. Use API Wallet credit, or an eligible Builder subscription key. - Copy an exact model ID from the catalog returned for this key. Replace MODEL_ID_FROM_CATALOG wherever it appears below. - Use a separate named key for this app. Paste it only into the app's credential field; keep it out of chat prompts, screenshots, shared configuration and workflow exports. - AnythingLLM Desktop, or administrator access to your AnythingLLM installation. Have the selected model's context and output limits available. ## Setup 1. **Open the language-model settings** Open Settings → AI Providers → LLM and choose Generic OpenAI. During first-time setup, this selection appears under LLM Preference. These settings establish the system default for workspaces without an override. 2. **Enter your connection** Set Base URL to https://api.lumaos.cloud/v1 and paste your Luma key into API Key. In Selected Model, choose the catalog ID you copied. If discovery is unavailable and the field permits manual input, enter that ID directly. 3. **Set the token limits** Model context window is the total space available to the conversation. Max Tokens is the output budget for a reply. Use the selected model's published limits and keep Max Tokens within its output limit; do not reuse another model's values. 4. **Save and check the workspace** Click Save changes. Open the intended workspace's settings → Chat Settings and check Workspace LLM Provider and Workspace Chat model. An existing workspace override takes precedence over the system default. 5. **Try a plain conversation** Use Chat mode for the first test, open a fresh conversation in that workspace, and send “Reply with hello.” Agent Configuration is separate; configure it later if you need agent tools. ## Connection fields - LLM Provider: Generic OpenAI - Base URL: https://api.lumaos.cloud/v1 - API Key: Your Luma Cloud API key - Selected Model: MODEL_ID_FROM_CATALOG - Model context window: The selected model's published context limit - Max Tokens: An output budget within that model's limit ## Verify - Wait for the workspace to show a completed text answer. Query mode can require document context, so use Chat for this test. - After the reply, open Dashboard → Usage. Check API Wallet for a wallet key, or subscription usage for a Builder subscription key. - If the selected model is wrong, inspect Workspace Chat model and the separate Agent Configuration before changing your Luma key. ## Limits - Configure document embeddings separately. Selecting this chat provider does not make Luma Cloud an embedding service. - Agent tools depend on the model and your AnythingLLM version. Check a simple tool action before relying on an unattended workflow. - A shared installation's system model may serve several workspaces. Use a key intended for that installation and check workspace access before adding personal credentials. ## Troubleshooting ### 401 or an invalid-key message Confirm that the saved credential contains a Luma Cloud API key, without surrounding quotes, spaces or a Bearer prefix. Check that this key is still active in Dashboard → API. A dashboard password or another provider's key will not work. ### The workspace still answers with another model Open that workspace's Chat Settings. Its provider and model can override the system default. For @agent conversations, inspect Agent Configuration separately, then start a fresh conversation. ### Selected Model is empty Recheck Base URL https://api.lumaos.cloud/v1 and the key. Current Generic OpenAI settings offer manual model entry when discovery returns no models. Enter an exact available catalog ID; do not use a display name or keep the placeholder. ### Documents cannot be embedded or Query returns no answer Test Chat mode without documents first. Configure the embedding provider separately and finish document processing before Query mode. A successful LLM reply does not test document indexing. ### Context or output limit error Use the selected model's limits for Model context window and Max Tokens. Start a new short conversation and reduce retrieved document context. Increasing the local setting cannot increase the model's real limit. ### 429, a balance warning or an interrupted run Read the error and check the usage associated with this key. A wallet key uses API Wallet funds; a Builder subscription key uses its eligible subscription quota. Pause retries, wait for the indicated reset or retry time, and reduce parallel requests. After a timeout, check the existing result before running a costly job again. ## Official client documentation - [AnythingLLM: OpenAI Generic](https://docs.anythingllm.com/setup/llm-configuration/cloud/openai-generic) - [AnythingLLM: system, workspace and agent models](https://docs.anythingllm.com/setup/llm-configuration/overview) - [AnythingLLM: provider settings source](https://github.com/Mintplex-Labs/anything-llm/blob/master/frontend/src/components/LLMSelection/GenericOpenAiOptions/index.jsx) - [AnythingLLM: settings navigation](https://github.com/Mintplex-Labs/anything-llm/blob/master/frontend/src/components/SettingsSidebar/index.jsx) - [AnythingLLM: interface labels and chat modes](https://github.com/Mintplex-Labs/anything-llm/blob/master/frontend/src/locales/en/common.js)