Developer documentation
Build with Eldrim.
Bring your code. We’ll mind the machinery. Connect your application, use open models in your coding tools, and schedule work from your own system.
Updated 4 October 2026. Early access: confirm model access, capacity and service terms during onboarding.
Make your first request
You need an Eldrim account, an API key and a positive evaluation balance. Request access if you are new, or use the console to check an existing account. An OpenAI or Anthropic key will not authenticate with Eldrim.
The API base URL is https://api.eldrim.net/v1. Keep the key on your server or development machine, outside source control. Browser applications should call your backend rather than expose the key in frontend code.
1. List the models your account can use
# Bash: enter your Eldrim key without saving it in shell history.
read -rsp 'Eldrim API key: ' ELDRIM_API_KEY; echo
export ELDRIM_API_KEY
curl --fail-with-body https://api.eldrim.net/v1/models \
-H "Authorization: Bearer $ELDRIM_API_KEY"
The response contains a data array with model IDs, aliases, prices, a primary endpoint and an endpoints array of supported protocols for your account. Choose a model whose endpoint is /v1/chat/completions for this example. Access and availability vary by account.
If data is empty, your key authenticated but no model is enabled for your account and residency policy. The response includes eldrim.code: no_models_available. Contact support@eldrim.ai to review access. Creating a key does not automatically grant access to restricted test hardware.
2. Send a Chat Completions request
# Use a chat model returned by /v1/models. This is an example id.
export ELDRIM_MODEL=qwen3.6-35b
curl --fail-with-body --max-time 240 -i \
https://api.eldrim.net/v1/chat/completions \
-H "Authorization: Bearer $ELDRIM_API_KEY" \
-H "Content-Type: application/json" \
-d "{\"model\": \"$ELDRIM_MODEL\",
\"messages\": [{\"role\": \"user\", \"content\": \"Explain a queue in one sentence.\"}],
\"max_tokens\": 256}"
messages is the conversation: a list of objects with a role and content. Use system for instructions, user for input and assistant for earlier answers. Send the relevant history with each request. max_tokens bounds generated output; the prompt and output must also fit the model’s active context window.
3. Read the answer and token usage
This is an illustrative response, not a measured token count or a promised model answer:
{
"id": "chatcmpl-example",
"object": "chat.completion",
"model": "qwen3.6-35b",
"choices": [{
"index": 0,
"message": {"role": "assistant", "content": "A queue holds work until a worker can process it."},
"finish_reason": "stop"
}],
"usage": {"prompt_tokens": 24, "completion_tokens": 16, "total_tokens": 40}
}
The answer is in choices[0].message.content. Check finish_reason: length means an output limit was reached. Tool-capable models may return tool_calls instead of a text answer. Your application validates and executes those calls, then sends their results back.
The API at a glance
Read the full API reference or download the OpenAPI 3.1.1 specification. The same machine-readable contract is available at api.eldrim.net/openapi.json. Import it into Swagger-compatible tools. The reference includes schemas, authentication, examples, responses and limitations.
Eldrim implements a subset of the OpenAI API shapes. Your client uses Eldrim’s URL, key and model IDs. An SDK feature being available in OpenAI’s documentation does not mean Eldrim serves it.
| Endpoint | Use | Current scope |
|---|---|---|
GET /v1/models | Discover models and prices | Filtered by account visibility and residency. |
GET /v1/models/{model} | Inspect one model or alias | Same access rules as the model list; includes supported endpoints. |
POST /v1/completions | Complete a text prompt | One string prompt; JSON or streaming on enabled routes. |
POST /v1/embeddings | Create text vectors | One string or up to 128 strings; embedding models only. |
POST /v1/chat/completions | Chat and application inference | JSON responses or streaming; model-dependent tools. |
POST /v1/responses | Codex and Responses clients | Stateless requests on enabled routes; client-executed function tools. |
POST /v1/messages | Claude Code and Messages clients | Native Messages format on enabled routes; client-executed tools. |
GET /v1/eldrim/account | Balance and usage | Your account and usage totals for the last 30 days. |
GET /v1/eldrim/requestsGET /v1/eldrim/requests/{id} | Inspect execution records | Your account’s request metadata, route and cost; no prompt or response content. |
POST /v1/eldrim/topup | Open a checkout | Stripe test mode during evaluation; integer amount_sek. |
POST /v1/messages/count_tokens | Count tokens | Capability gate implemented; no enabled provider route today. Returns an explicit error. |
POST /v1/systemone | Classification and scoring | Decision models with a separate request format. Confirm suitability with Eldrim. |
Authenticate with Authorization: Bearer YOUR_ELDRIM_KEY. Messages also accepts x-api-key; a supplied bearer header takes precedence. The unauthenticated GET /healthz endpoint checks gateway health, not model readiness or available GPU capacity.
Hosted Batch jobs, files/uploads, vector stores, fine tuning, media generation, moderation and realtime sessions are unsupported and return JSON errors. Store vectors and schedule jobs in your application. Responses does not provide stored conversations, previous_response_id, background jobs or provider-hosted tools. /v1/messages/count_tokens has no enabled route today. Use inline content and send conversation history from your client.
Use the OpenAI Python SDK
Install the SDK in your project’s Python environment:
python3 -m pip install openai
The key and model variables are the ones set in the first example. This request uses Chat Completions:
import os
from openai import OpenAI
client = OpenAI(
base_url="https://api.eldrim.net/v1",
api_key=os.environ["ELDRIM_API_KEY"],
timeout=240.0,
max_retries=0, # Choose retry policy explicitly for your workload.
)
reply = client.chat.completions.create(
model=os.environ["ELDRIM_MODEL"],
messages=[{"role": "user", "content": "Explain a queue in one sentence."}],
max_tokens=256,
)
print(reply.choices[0].message.content)
print(reply.usage)
For a Responses-capable model and route, use the same client with the Responses format:
reply = client.responses.create(
model=os.environ["ELDRIM_MODEL"],
input="Explain a queue in one sentence.",
max_output_tokens=256,
store=False,
)
print(reply.output_text)
input, max_output_tokens and output_text belong to Responses; messages, max_tokens and choices belong to Chat Completions. The gateway selects only routes that support the requested protocol.
OpenAI’s Chat Completions reference describes the upstream format. Use the Eldrim support table above when choosing features.
Stream an answer
Set stream=True to receive text as it is generated. Reuse the Python client above. The usage-only chunk may have an empty choices list:
with client.chat.completions.create(
model=os.environ["ELDRIM_MODEL"],
messages=[{"role": "user", "content": "Explain a queue in one sentence."}],
max_tokens=256,
stream=True,
stream_options={"include_usage": True},
) as stream:
for chunk in stream:
if chunk.choices:
print(chunk.choices[0].delta.content or "", end="", flush=True)
if chunk.usage is not None:
print("\nUsage:", chunk.usage)
On the wire, Chat Completions uses server-sent events (SSE) and normally ends with data: [DONE]. Responses completes with response.completed; Messages with message_stop. Treat a missing terminal event or an SDK interruption as incomplete output, even if the HTTP status began as 200.
Once a stream starts, Eldrim keeps it on that provider. A broken answer is never continued by another provider. Closing the connection cancels the upstream request; reported usage may still be charged. Keep partial output separate from completed results.
Choose a model for the job
Start with /v1/models and a small example of your actual task. The list describes what your account may use, not which models are warm or how much concurrent capacity is free. Use endpoints to check the protocols allowed for that model under your account’s policy. The primary endpoint field identifies its main task. Protocol support does not guarantee that every model accepts every optional field.
- Chat and extraction: compare output quality, latency and token use on representative inputs.
- Coding: check tool calling, instruction following and multi-step edits, including whether your tests pass.
- Long inputs: confirm the server’s active context window. A model’s advertised maximum does not establish the deployed budget.
- Decisions: use the dedicated decision endpoint only with a suitable model and reviewed use case.
qwen3.6-35b passed small Codex and Claude Code tool round trips on 1 October 2026. That establishes basic connectivity and tool use. It does not establish parity with proprietary models or reliability across a large codebase. Confirm access before selecting it.
Complete a text prompt
For clients using the older Completions format, choose a model whose endpoints includes /v1/completions. The gateway forwards the native protocol to an approved provider. It accepts one string prompt and one completion per request; prompt arrays, token IDs, n > 1 and best_of > 1 are refused.
reply = client.completions.create(
model="llama3.2-3b", # Confirm access in /v1/models first.
prompt="The capital of Sweden is",
max_tokens=16,
)
print(reply.choices[0].text)
With stream=True, read choices[0].text from each non-empty chunk. Set stream_options={"include_usage": True} for the final usage chunk. The stream ends with [DONE]. Provider-dependent options such as suffix, echo and log probabilities must be tested with the selected model.
Turn text into vectors
Embeddings support search, retrieval and similarity comparisons. Select a model with /v1/embeddings in its endpoints list. Authorised test accounts can use nomic-embed-text on the existing Swedish test service. Its default output has 768 dimensions. This is evaluation capacity; it is not contracted production capacity.
vectors = client.embeddings.create(
model="nomic-embed-text",
input=[
"search_document: Recovered heat can warm a building.",
"search_query: district heating",
],
encoding_format="float",
)
print(len(vectors.data[0].embedding))
print(vectors.usage.prompt_tokens)
Send one non-empty string or a batch of up to 128 non-empty strings. Token arrays and streaming are unsupported. The endpoint accepts float or base64 encoding; optional reduced dimensions depend on the model and provider. Only input tokens are charged, using the model’s integer input rate.
For Nomic retrieval, use search_document: for indexed text and search_query: for queries, following the model’s task-prefix guidance. Store the returned vectors in your own index. Eldrim does not provide a vector store or a file upload service.
Check the route against your policy
Residency is configured on your account during onboarding. A Sweden-only policy permits Swedish routes; an EU policy permits Swedish and other EU routes. There is no per-request residency override. Requests and failover stay within the configured policy.
When a provider answers, the response includes these headers. The values below are illustrative:
x-request-id: req_example
x-eldrim-provider: configured-provider
x-eldrim-residency: se
Save the request ID with your job record for support and reconciliation. Provider and residency headers identify the configured route. They are not independent hardware-location attestation. Data handling, provider terms and regulated uses still need their own review. Read about Eldrim’s sovereignty approach.
Run Codex CLI on Eldrim
The coding harness stays on your machine; model inference runs through Eldrim. Codex needs the Responses API. An OpenAI-compatible Chat Completions endpoint alone is insufficient.
- Install Codex CLI and set
ELDRIM_API_KEYas above. - Download the Eldrim profile, inspect it and save it as
~/.codex/eldrim.config.toml. Preserve an existing file with that name. - Confirm that the model’s
endpointsincludes/v1/responses, then start Codex:
codex --profile eldrim
View the complete Codex profile
# Copy to ~/.codex/eldrim.config.toml, then: codex --profile eldrim
# Keep ELDRIM_API_KEY in the environment. This is a separate opt-in profile.
model = "qwen3.6-35b"
model_provider = "eldrim"
web_search = "disabled"
# Confirm the server's active context before raising these limits.
model_context_window = 32768
model_auto_compact_token_limit = 24000
model_reasoning_effort = "none"
[model_providers.eldrim]
name = "Eldrim"
base_url = "https://api.eldrim.net/v1"
env_key = "ELDRIM_API_KEY"
wire_api = "responses"
supports_websockets = false
request_max_retries = 0
stream_max_retries = 0
stream_idle_timeout_ms = 240000
[analytics]
enabled = false
[features]
# The native subset accepts function tools, not tool namespaces.
multi_agent = falseThis opt-in profile leaves other Codex sessions on their usual provider. It uses Responses over HTTP/SSE, disables hosted web search and namespaced multi-agent tools, and leaves normal sandbox and approval controls in place. Client retries are disabled so your workload can choose its own retry policy.
The profile starts at a 32,768-token context budget with earlier compaction. Confirm that budget with Eldrim before a large run. Custom-model metadata warnings are currently expected; client defaults are not a capacity guarantee. Official Codex configuration guidance.
Run Claude Code on Eldrim
Claude Code can use Eldrim’s Messages endpoint to run an open model. This does not provide Anthropic’s Claude model. Ollama documents this integration; Anthropic does not officially support non-Claude models.
- Install Claude Code and set
ELDRIM_API_KEY. - Download and inspect the Bash launcher. It sets the base URL to
https://api.eldrim.net, without/v1, and pins model aliases to the chosen Eldrim model. - Run it in your project, then use
/statusto confirm the effective URL and credential source:
# Save the launcher in your working directory and inspect it first.
bash ./claude-eldrim
View the complete Claude Code launcher
#!/usr/bin/env bash
# One-session configuration. Does not write global Claude settings.
set -euo pipefail
: "${ELDRIM_API_KEY:?Set ELDRIM_API_KEY to your Eldrim key}"
eldrim_model="${ELDRIM_MODEL:-qwen3.6-35b}"
export ANTHROPIC_BASE_URL=https://api.eldrim.net
export ANTHROPIC_AUTH_TOKEN="$ELDRIM_API_KEY"
export ANTHROPIC_API_KEY=""
export ANTHROPIC_MODEL="$eldrim_model"
export ANTHROPIC_DEFAULT_OPUS_MODEL="$eldrim_model"
export ANTHROPIC_DEFAULT_SONNET_MODEL="$eldrim_model"
export ANTHROPIC_DEFAULT_HAIKU_MODEL="$eldrim_model"
export ANTHROPIC_DEFAULT_FABLE_MODEL="$eldrim_model"
export CLAUDE_CODE_SUBAGENT_MODEL="$eldrim_model"
export CLAUDE_CODE_DISABLE_NONESSENTIAL_TRAFFIC=1
export DISABLE_PROMPT_CACHING=1
# Prevent Claude-specific adaptive thinking/context-management requests.
# The open model may still reason according to its own server settings.
export MAX_THINKING_TOKENS=0
# Keep the normal permission flow. Existing managed settings still apply.
exec claude --model "$eldrim_model" --disallowedTools WebSearch WebFetch "$@"Set ELDRIM_MODEL to choose another supported model. The launcher disables built-in web tools, nonessential traffic and Claude-specific adaptive context-management requests. Normal file and command permissions remain active. Existing managed settings can override environment variables; background or cloud agents need separate configuration.
Codex CLI 0.157.0 and Claude Code 2.1.282 were checked with a real file-reading tool call and a follow-up answer. Recheck after client or model upgrades. Client cost estimates and context displays may use vendor defaults; Eldrim’s ledger and agreed capacity are the relevant values.
Schedule workloads from your own system
Today, you own the schedule and queue. Start work from cron, a systemd timer, CI or your application’s worker queue. Eldrim serves the inference requests. There is no schedule_at parameter, hosted job queue, reservation API or guaranteed off-peak price.
- Your timer
- Your job file or queue
- Bounded worker
- Eldrim API
- Your saved result
Start with one in-flight request. Measure latency and output quality, then agree capacity before increasing concurrency. A burst of requests does not reserve a GPU. For a fixed deadline or sustained load, discuss volume, token sizes, completion window and residency with Eldrim.
A resumable local batch
Download and inspect the Python batch example. It uses Python 3.10+ with no extra packages, processes one job at a time, and saves each response with its request ID and route headers. Create jobs.jsonl:
{"id":"queue-001","prompt":"Explain a queue in one sentence."}
{"id":"queue-002","prompt":"Explain backpressure in one sentence."}
# Python 3.10+. Uses ELDRIM_API_KEY and ELDRIM_MODEL from your environment.
python3 eldrim-batch.py jobs.jsonl --output-dir results --max-tokens 256
Results go to results/queue-001.json and results/queue-002.json. A rerun skips saved jobs. Changing a saved job’s prompt, model or output limit requires a new job ID. The script stops on a failed, truncated or tool-call response, leaving earlier results intact. Review the output limit if generation stops at its token budget. Keep the input and result files private; they contain your workload data.
The default is one attempt. --attempts 3 opts into bounded retries for transient HTTP or transport failures with exponential backoff and jitter. A lost connection can leave completion uncertain. A retry or a restart after an unsaved response may run inference again and incur another charge. Local checkpoints do not provide exactly-once execution.
Run it each night with Linux cron
Place the script and job file in a directory you own, for example /srv/eldrim-jobs. Create eldrim.env there with your real values and restrict it with chmod 600 eldrim.env. Keep it out of Git:
ELDRIM_API_KEY=replace-with-your-eldrim-key
ELDRIM_MODEL=replace-with-an-account-visible-chat-model
Save this wrapper as run.sh. It gives cron an explicit working directory and loads the variables that an interactive shell would normally supply:
#!/bin/sh
set -eu
umask 077
cd /srv/eldrim-jobs
set -a
. ./eldrim.env
set +a
exec /usr/bin/python3 ./eldrim-batch.py ./jobs.jsonl \
--output-dir ./results --max-tokens 256 --timeout 180
Add this entry with crontab -e, after confirming the paths to python3, flock and GNU timeout on your Linux host:
# Daily at 02:15 in this machine's cron timezone.
15 2 * * * cd /srv/eldrim-jobs && /usr/bin/flock -n ./batch.lock /usr/bin/timeout --kill-after=10s 45m /bin/sh ./run.sh >>./batch.log 2>&1
The lock prevents overlapping runs on this host. The timeout bounds the whole run; the Python socket timeout alone does not provide a wall-clock deadline. If the previous job still holds the lock, that scheduled run is skipped. Monitor failures and skipped runs, rotate batch.log, and use a durable queue when missed schedules need catching up.
The cron expression follows the machine’s timezone, including its daylight-saving behaviour. Use a UTC-configured scheduler for a fixed daily cadence. Add new job IDs for new daily work; rerunning the same file will skip completed jobs.
Track usage and set a budget
curl --fail-with-body https://api.eldrim.net/v1/eldrim/account \
-H "Authorization: Bearer $ELDRIM_API_KEY"
The account response includes residency, balance_sek, integer balance_micro_sek, and usage_30d totals per model. /v1/models quotes input and output prices in öre per million tokens. One SEK is 100 öre or 1,000,000 micro-SEK. Evaluation pricing and production terms must be confirmed before relying on a cost estimate.
Use GET /v1/eldrim/requests?limit=20 to see your most recent execution records, or GET /v1/eldrim/requests/{id} for a request ID. Records include provider, residency, status, routing and integer cost charged to your account. They do not contain prompts or answers. Requests refused before routing may have no record.
Create a key per worker or environment in the console, so each one's requests can be told apart by key_id and revoked on its own. A key can expire after 30 to 365 days, chosen when it is created. Per-key spending budgets are not implemented yet. A request is refused when the balance is exhausted. Already-running requests can take it below zero, so the balance is not a strict spending cap. Bound concurrency, input size and generated tokens in your own worker. Record usage and check the account between batches.
Eldrim is a small team, founded by Erik Stenman and operated by Happi Hacking AB. Meet the people and find our contact details.
For support, send the request ID, model, approximate time and error status to support@eldrim.ai. Keep API keys and confidential prompts out of routine logs and support messages.
Handle failures deliberately
{"error":{"type":"insufficient_quota","message":"Balance exhausted. Top up to continue."}}
| Status | Check | Next action |
|---|---|---|
| 400 | Request shape, model endpoint or unsupported protocol feature | Correct the request. Repeating it unchanged will not help. |
| 401 | Missing, invalid or revoked Eldrim key | Check the key and where the client sends it. |
| 402 | Evaluation balance exhausted | Check the console or contact Eldrim before resuming. |
| 403 | Suspended account or no route within policy | Resolve access or capacity with Eldrim. |
| 404 | model_not_found, missing request record or endpoint_not_found | Read the error code. Check model access or the exact URL against the reference. |
| 405 | method_not_allowed | Use the HTTP method in the Allow header. |
| 501 | endpoint_not_supported | This API family is not implemented. Follow the documentation link; retrying will not help. |
| 413 | Request too large | Reduce or split the input, preserving the task’s context. |
| 429 / 502 / 503 / 504 | Capacity, upstream failure or transport interruption | Back off, cap attempts and reduce concurrency. A retry may be charged again. |
Gateway JSON errors carry an X-Request-Id. Unsupported paths return a JSON error instead of an empty body. HEAD requests deliberately contain headers only; OPTIONS on supported JSON endpoints returns 204 with an Allow header. Use curl -i to inspect the status and headers when diagnosing a response.
The gateway may already have tried other eligible routes before returning 502. Authentication and policy errors can be returned before a provider is selected, so route headers may be absent. Errors after streaming starts cannot replace its initial HTTP 200; detect incomplete SSE and record the job as interrupted.
Chat, text completions, embeddings, Responses and Messages request bodies are limited to 8 MiB; streamed events to 2 MiB. The model’s context budget is a separate, usually tighter constraint. System One has a 128 KiB request limit. Use timeouts appropriate to model loading and generation, with a separate overall job deadline.