01
Quickstart#
Three steps to your first answer.
- Sign in with your email. For API credit, see Credits.
- Create a key under API keys. It is shown once; we store only a hash of it. Keep it in the
PERYN_API_KEYenvironment variable. - Point any OpenAI SDK at the base URL and send a request.
curl https:// api. peryn. ai/ v1/ chat/ completions \-H "Authorization: Bearer $PERYN_ API_ KEY" \-H "Content-Type: application/ json" \-d '{"model": "llama-3.1-8b-instruct","max_tokens": 64,"messages": [{"role": "user", "content": "Say hello in five words."}]}'The answer has the OpenAI shape, with the token counts you pay for in usage:
{"id": "chatcmpl-…","object": "chat. completion","created": 1791216000,"model": "llama-3.1-8b-instruct","choices": [{"index": 0,"message": { "role": "assistant", "content": "Hello! Nice to meet you." },"logprobs": null,"finish_ reason": "stop"}],"usage": { "prompt_ tokens": 16, "completion_ tokens": 8, "total_ tokens": 24 }}Your code
Any OpenAI SDK, our base URL, your API key.
Our gateway
Checks the key, the rate limits and your credit, then picks a GPU.
A GPU node
Runs the model. On Standard, the same work also mines PRL.
02
Base URL and models#
One base URL for every endpoint. Send your key as a Bearer token.
https:// api. peryn. ai/ v1Code that still uses https://api.ai.pearlsafe.xyz/v1 keeps working: it is the same API.
Every request that runs a model needs the header Authorization: Bearer <key>. Production keys start with pc_live_. Bodies are JSON, up to 2 MB.
https://api.peryn.ai/v1
- POST
/chat/completionsChat completions, streamed or not.API key - POST
/completionsText completions from one prompt string.API key - GET
/modelsEvery model: id, context, price.No key - GET
/models/{id}One model.No key - GET
/catalogEvery model with its price per million tokens, context length and licence.No key - GET
/catalog/{id}One model from the catalog.No key
curl https:// api. peryn. ai/ v1/ modelsRequest fields
The fields below are read; other OpenAI fields that do not change the answer (such as user or metadata) are ignored.
model- A model id from the table below. Required.
messages- Chat only. 1 to 2,048 messages; roles system, developer, user and assistant; text content (a string or text parts).
prompt- Text completions only. One string.
max_tokens- Output limit, 1 up to the model’s max output. max_completion_tokens works the same and wins if both are set.
stream- true for server-sent events.
stream_options- {"include_usage": true} adds a last chunk with token counts.
temperature- 0 to 2.
top_p- Above 0, up to 1.
top_k- An integer, -1 or more.
min_p- 0 to 1.
frequency_penalty, presence_penalty- -2 to 2.
repetition_penalty- Above 0, up to 2.
seed- An integer.
stop- A string or up to 4 strings, 256 characters each.
logit_bias- Token id to bias, -100 to 100.
response_format- Chat only. type text, json_object or json_schema.
n- 1.
Models take text in and give text out. Tool and function calling, images, logprobs, echo, suffix and n above 1 are answered with a 400 that names the field. If you need one of them, talk to us.
Text completions
curl https:// api. peryn. ai/ v1/ completions \-H "Authorization: Bearer $PERYN_ API_ KEY" \-H "Content-Type: application/ json" \-d '{"model": "llama-3.1-8b-instruct", "prompt": "The capital of France is", "max_tokens": 16}'Response headers
x-request-idon every response. Quote it when you write to us about a request.x-pc-node-countryon every response a GPU node served, streamed or not: the node’s country as a two-letter code (ISO 3166-1).
curl -sS -D - -o /dev/null https:// api. peryn. ai/ v1/ chat/ completions \-H "Authorization: Bearer $PERYN_ API_ KEY" \-H "Content-Type: application/ json" \-d '{"model": "llama-3.1-8b-instruct", "max_tokens": 8, "messages": [{"role": "user", "content": "Hi"}]}'03
Streaming#
Server-sent events in the OpenAI format: tokens arrive as the model writes them.
Set stream: true. Each event is a data: line with a chat.completion.chunk; the text is in choices[0].delta.content and a final chunk carries finish_reason. With stream_options: {"include_usage": true} every chunk has "usage": null and one more chunk, with empty choices, carries the token counts. The stream ends with data: [DONE].
curl -N https:// api. peryn. ai/ v1/ chat/ completions \-H "Authorization: Bearer $PERYN_ API_ KEY" \-H "Content-Type: application/ json" \-d '{"model": "llama-3.1-8b-instruct","max_tokens": 128,"stream": true,"stream_ options": {"include_ usage": true},"messages": [{"role": "user", "content": "Write a haiku about GPUs."}]}'On the wire (abridged)
: ping
data: {"id": "chatcmpl-…", "object": "chat. completion. chunk", "created": 1791216000, "model": "llama-3.1-8b-instruct", "choices": [{"index": 0, "delta": {"role": "assistant", "content": "Silicon"}, "logprobs": null, "finish_reason": null}], "usage": null}
data: {…, "choices": [{"index": 0, "delta": {"content": " hums"}, "logprobs": null, "finish_reason": null}], "usage": null}
data: {…, "choices": [{"index": 0, "delta": {}, "logprobs": null, "finish_reason": "stop"}], "usage": null}
data: {…, "choices": [], "usage": {"prompt_tokens": 16, "completion_tokens": 19, "total_tokens": 35}}
data: [DONE]- Lines that start with a colon (
: ping) keep the connection open while the model works on its first token. The OpenAI SDKs skip them. - An error before the first event comes back as a normal HTTP error. An error after the stream has started arrives as an
event: errorline with the error JSON, thendata: [DONE]. - If the stream breaks or you close it, you pay for the tokens that reached you and nothing more. If none did, you pay nothing.
- A request that waits for a GPU sends nothing until a GPU takes it (see Models and prices).
04
Models and prices#
Live prices in US dollars per million tokens, excluding VAT.
Put the model id in the model field. These are Standard prices. Each follows the market: 15% below the cheapest comparable listed provider. It never falls under what we pay the GPU owner plus our margin.
llama-3.1-8b-instructAvailableLlama 3.1 8B Instruct · 8B · 32,768 tokens, max output 8,192
- Input
- $0.017
- Output
- $0.034
The same table is JSON at https://app.peryn.ai/api/pricing and, with context lengths, at GET /v1/catalog. The comparison with the market is on the pricing page; every model’s licence is on the licences page.
curl https:// api. peryn. ai/ v1/ catalogWhat a request costs
Input tokens times the input price plus output tokens times the output price, rounded up to the next millionth of a dollar. We count the tokens ourselves with the model’s own tokenizer; the counts are the usage in every answer and on your usage page.
05
Privacy tiers#
New keys use the Standard tier. Private is on request: talk to us.
Standard
No training on prompts or completions, ever. Zero retention: we do not log or store prompt or completion content; we keep only request metadata (such as request ID, account, timestamps, model, token counts, latency, cost and status) for billing and reliability. The same GPU work also mines PRL; Mining and your data in the privacy policy explains what that means. The prices above are Standard prices.
Where your requests run
By default, a request runs on any healthy node, in any country. Set a region rule for your whole account under Settings. Narrow it for one key on the API keys page. A rule is EU only, US only, US + EU, your own list of countries, or countries to exclude. Requests then run only on nodes in the allowed countries. If none of them takes it within the wait, the request fails with 503 and code no_capacity_in_region. It never runs elsewhere instead.
A node’s country comes from its operator’s declaration, checked against the location of its IP address. A dishonest operator could fake both (for example with a VPN); we remove nodes we catch. Each answer names its node’s country in x-pc-node-country, and your usage page breaks your requests down by country.
06
Rate limits and credits#
Per-key limits per minute, and prepaid credit in US dollars.
Prepaid credit
- Held while a request runs: its worst case
- Charged when it ends: the tokens it used
- The rest of the hold goes back to your credit
Limits per key
- 60 requests per minute
- 200,000 tokens per minute
Rate limits
- Each key allows 60 requests and 200,000 tokens per minute. Both refill continuously. The API keys page shows each key’s limits. Need more? Talk to us.
- A request counts its input tokens plus
max_tokensagainst the token limit when it starts. When it ends, it gets back what it didn’t use. Withoutmax_tokensit counts the model’s full output allowance (up to 8,192 tokens), so setting it lets more requests run. - Over a limit, the request is refused at once with
429, coderate_limit_exceeded, and aRetry-Afterheader in seconds. Nothing is queued; wait that long and retry.
Credits
- Credit is prepaid in US dollars, excluding VAT. Card checkout opens at launch. To get API credit now, talk to us.
- When a request starts we hold its worst case: input tokens plus
max_tokensat the model’s price. When it ends we charge what it used and release the rest. If your available credit can’t cover the hold, the request is refused with402, codeinsufficient_quota. - A request that fails before any token reaches you costs nothing.
- The usage page shows requests, tokens and cost per key, model and country, for the last 7, 30 or 90 days.
- Unused credit expires after 12 months without activity (a sign-in, a request or a purchase). We email you at least 30 days before.
07
Errors#
Errors use the OpenAI shape, so the OpenAI SDKs raise them as exceptions with the status code.
Error body
HTTP/1.1 404 Not Found
{ "error": { "message": "The model 'no-such-model' does not exist or is not available.",Written for people. Log it or show it. "type": "invalid_request_error",The kind of error. "param": "model",The field at fault, when there is one. "code": "model_not_found"Stable. Branch on it. }}Errors that can clear by themselves (429, 503) carry a Retry-After header.
- 400
invalid_request_error invalid_jsoninvalid_valuecontext_length_exceededunsupported_parameterunsupported_valueunsupported_contentbad_requestThe request needs fixing. The body is not JSON, a field is out of range, or prompt plus max_tokens is over the context. Other causes: a feature the API does not take, or input the model refused. param names the field. Retrying unchanged gives the same answer.- 401
authentication_error missing_api_keyinvalid_api_keyrevoked_api_keyNo key, a key we do not know (or a malformed Authorization header), or a revoked key.- 402
insufficient_quota insufficient_quotaYour available credit does not cover this request’s hold. Lower max_tokens, or get more credit (see Rate limits and credits).- 403
permission_error account_disabledThe account is disabled. Write to us.- 404
invalid_request_error model_not_foundnot_foundUnknown model id, or a path the API does not have.- 413
invalid_request_error request_too_largeThe body is larger than 2 MB.- 429
rate_limit_error rate_limit_exceededno_capacityrate_limit_exceeded: your key’s requests or tokens per minute are used up. no_capacity: every GPU for this model stayed busy. Wait for Retry-After, then retry.- 500
server_error internal_errorregion_rule_invalidSomething failed on our side. Retry with backoff. region_rule_invalid: save your region rule again in the console.- 502
server_error upstream_errorThe GPU node serving the request failed. Retry.- 503
server_error no_capacityno_capacity_in_regionNo GPU could take the request within the wait (see Models and prices), or none in the countries your region rule allows. Retry after Retry-After.- 504
server_error timeoutThe model did not finish within 5 minutes. Retry, or lower max_tokens.
curl -i https:// api. peryn. ai/ v1/ chat/ completions \-H "Authorization: Bearer $PERYN_ API_ KEY" \-H "Content-Type: application/ json" \-d '{"model": "no-such-model", "messages": [{"role": "user", "content": "Hi"}]}'Retry 429, 500, 502, 503 and 504 with backoff; the OpenAI SDKs do this for you (twice by default). Fix the request for the other codes. Stuck? Write to hello@peryn.ai with the x-request-id.
08
For GPU owners#
Your GPU can serve these requests and mine PRL with the same work.
Your GPU mines PRL through our pool with the same work that serves requests. Payouts begin at launch. Want paid jobs from our API too? Tell us about your fleet. Our miner image runs the model and our node software, which never logs or stores prompts or completions.
What you need
- An NVIDIA H100 or H200 (x86_64).
- NVIDIA driver r580 or newer (CUDA 13).
- A container with 16 GB of shared memory, and disk for our image and the model weights.
- Outbound HTTPS and WebSocket only. No inbound ports: the node dials out to us.
Connect a node
- Sign in to the miner portal and choose Add a node. Give it a name, its country and a model. You get an enrolment token: single use, valid for 24 hours.
- Run the command the portal shows, with the image reference we send you, on RunPod, Vast.ai or your own server. The node enrols and connects to our gateway and our pool. Ask us for the image reference: tell us about your fleet.
- The node mines from the start and shows as pending. Customer requests start once we approve it: we check who you are, your signed operator terms, its test requests and its country.
To mine without serving API requests, tick Pool only (no API traffic) when you create the token. The node then earns PRL only and needs no approval.
The pool keeps a 1% fee on block rewards, and a block reward counts after 100 confirmations. Earnings, estimates and payouts are on the GPU clouds page.