Skip to main content
Your GGUF Cloud deployment exposes the standard OpenAI-compatible routes, backed by your model’s llama-server. Point any OpenAI SDK or compatible client at your deployment’s base URL with /v1 appended.

Endpoints

All paths are relative to your deployment base URL:

Request

Authenticate with your ModelsLab API key (see Authentication).

Body

Body Attributes

string
default:"local"
The model to use. A deployment serves a single model, so this can be "local" or the model id you deployed — either way the request is routed to your deployment’s model.
array
required
Array of message objects, each with a role (system, user, or assistant) and content. Used by /v1/chat/completions.
string
A single text prompt. Used by the /v1/completions endpoint instead of messages.
integer
default:"256"
Maximum number of tokens to generate. Range: 1 to the model’s context limit.
number
default:"0.8"
Sampling temperature. Lower values (0.10.3) produce focused, deterministic output; higher values increase creativity. Range: 0.02.0.
number
default:"0.95"
Nucleus sampling — only consider tokens with cumulative probability above this threshold. Range: 0.01.0. Use either temperature or top_p.
integer
Only sample from the top K most likely tokens (a llama.cpp sampling option).
number
default:"1.1"
Penalty applied to repeated tokens. Values > 1 discourage repetition (a llama.cpp sampling option).
number
default:"0"
Penalizes tokens that have already appeared, encouraging new topics. Range: -22.
number
default:"0"
Penalizes tokens proportionally to how often they have appeared. Range: -22.
string or array
One or more sequences where generation stops. The stop sequence is not included in the output.
integer
Seed for reproducible sampling. The same seed and input produce the same output.
boolean
default:"false"
When true, responses are streamed as Server-Sent Events (text/event-stream). See Streaming.

Response

Response Fields

string
Unique identifier for the completion.
string
The object type, e.g. chat.completion (or chat.completion.chunk while streaming).
string
The model that produced the response (your deployment’s model).
array
The generated choices. Each item contains an index, a message (with role and content), and a finish_reason.
object
Token accounting: prompt_tokens, completion_tokens, and total_tokens.

Streaming

Set "stream": true to receive Server-Sent Events (SSE) as tokens are generated:
Each SSE event carries a chat.completion.chunk object, terminated by data: [DONE]:

SDK Examples

This endpoint is a drop-in replacement for the OpenAI API. Just change the base_url and api_key:

Other OpenAI routes

/v1/embeddings is only available when the deployed model supports embeddings. If it does not, the endpoint returns an error from llama-server.