Inference API

Serverless inference at scale.

OpenAI and Anthropic SDK compatible. Deploy chat and coding model inference instantly. Pay per 1M tokens, scale automatically. No infrastructure to manage.

Latency
Sub-500ms response time
Pricing
Pay per token, no base fees
Models
Chat, coding, and more

Live model catalog.

Access our full roster of chat and coding models. Real pricing updated as rates change. OpenAI and Anthropic SDK compatible.

Base URL: https://api.runars.ca/v1

Aion Labs Aion 2.0🇨🇦
Input
$2.2784/M
Output
$4.5568/M
Context
128K tokens
Deepseek V4 Flash🇨🇦
Input
$0.2057/M
Output
$0.4114/M
Context
1M tokens
Deepseek V4 Flash 0731🇨🇦
Input
$0.1314/M
Output
$0.2645/M
Context
1M tokens
Deepseek V4 Pro🇨🇦
Input
$1.3313/M
Output
$3.9938/M
Context
1M tokens
Deepseek V4 Pro 0813🇨🇦
Input
$1.3313/M
Output
$3.9938/M
Context
1M tokens
GLM 4.5🇨🇦
Input
$0.9509/M
Output
$3.4866/M
Context
128K tokens
GLM 4.5 Air🇨🇦
Input
$0.3018/M
Output
$1.7433/M
Context
128K tokens
GLM 4.6🇨🇦
Input
$0.9509/M
Output
$3.4866/M
Context
198K tokens
GLM 4.7🇨🇦
Input
$0.9286/M
Output
$3.4866/M
Context
128K tokens
GLM 5🇨🇦
Input
$1.3930/M
Output
$4.4575/M
Context
198K tokens
GLM 5.1🇨🇦
Input
$2.2188/M
Output
$6.9733/M
Context
200K tokens
GLM 5.2🇨🇦
Input
$0.9509/M
Output
$3.0310/M
Context
1M tokens
GLM 5.3🇨🇦
Input
$2.2188/M
Output
$6.9733/M
Context
1M tokens
GLM 5.3 Flash🇨🇦
Input
$0.2089/M
Output
$0.6965/M
Context
1M tokens
Kimi K3🇨🇦
Input
$6.0512/M
Output
$30.2560/M
Context
1M tokens
Minimax M2.7🇨🇦
Input
$0.6051/M
Output
$2.4205/M
Context
198K tokens
Muse Spark 1.2🇨🇦
Input
$3.0616/M
Output
$10.4095/M
Context
1M tokens
Qwen 3.8 Max🇨🇦
Input
$5.0427/M
Output
$15.1280/M
Context
1M tokens
Qwen 3.5 35B A3B🇨🇦
Input
$0.6303/M
Output
$2.5213/M
Context
256K tokens
Qwen 3.6 35B A3B🇨🇦
Input
$0.2017/M
Output
$2.0171/M
Context
262K tokens
Qwen 3.6 27B🇨🇦
Input
$0.6965/M
Output
$4.6432/M
Context
256K tokens

Pricing shown in CAD. Billing by token. All prices subject to change — check docs for current rates.

Built for any workload.

Inference, training, batch jobs, agents. If it runs in a container, it runs on Runars.

3.0 Use Cases →

Inference

Run LLM inference endpoints that scale with demand.

Fine-tuning

Train and fine-tune models on your data, your schedule.

AI Agents

Long-running autonomous tasks without server management.

Batch Processing

Compute-heavy workloads that run when needed, scale to zero.

Ready to run AI in Canada?

Start with the inference API for instant compatibility. Scale to GPU rental when you're ready. No lock-in, no setup fees.

$npm install @anthropic-ai/sdk