Blazing Fast Agentic Inference
— One Endpoint

Run workloads on SOTA open-weight models, custom fine-tunes, with batch jobs — autoscaled, observable, failover resilient, hosted in India.

80

ms median

9,864

tok/sec

99.99

Performance, Control, Compliance.
Built for scale.

Built for teams to scale production-grade inference — within the SLA boundary.

01

.

Serverless & autoscale on demand. Failover built in. Zero cold-start concerns.

Concurrent requests · last 60s 39,988
-60s -45s -30s -15s now
Cold start
— ms
Autoscale
0
Failover
Multi-AZ

02

.

Scale to billions of tokens in hours.

No rate limits. No quotas. Ship the moment your workload spikes.

94ms
user model

03

.

Prompts and outputs never persist.

No logs. No storage. Nothing leaves memory after the response.

stdout
79920 d3e28 5f36a
/dev/null

04

.

Latency, throughput, cost, and failure rates — visible at every layer.

From the gateway to the GPU, every hop is instrumented and queryable.

Latency

Throughput

Cost

Failures

GPU

42 ms
68 ms
51 ms
94 ms
62 ms
112 ms
81 ms

SOTA models for text, image, and video

One API, every modality. Reason, generate, transcribe, and edit
— across the best open-weight models in each category, all
.running on the same elastic infra.

Namaste — serverless inference for the agentic
era
@registry-org

128K context

Reason, write, call tools.

Frontier open-source LLMs, tuned for lowest TTFT and highest throughput. Streaming, structured outputs, and native tool-calling — out of the box.

Kimi K2.6

GLM 5.1

DeepSeek V4 PRO

Flux · 1024×1024 · 2.1s

Generate, edit, upscale.

Flux, Qwen, and Stable Diffusion 3 — all running on dedicated image pods. Superfast generation with zero storage.

Flux Klein

Qwen-Image

Reason, write, call tools.

Frontier open-source LLMs, tuned for lowest TTFT and highest throughput. Streaming, structured outputs, and native tool-calling — out of the box.

Kimi K2.6

GLM 5.1

DeepSeek V4 PRO

Every open - weight model. One endpoint.

Hot-swap between Llama, Qwen, DeepSeek, Mistral, Gemma, and Sarvam. Bring your own checkpoint or deploy directly from Hugging Face with a single command.

Deepseek V4
deepseek-ai/DeepSeek-V4-Pro
40 tok/s
Kimi 2.6
moonshotai/Kimi-2.6
50 tok/s
GLM 5.1
THUDM/GLM-5.1
70 tok/s
Minimax 2.7
minimaxai/Minimax-2.7
120 tok/s
DeepSeek V4 Flash
deepseek-ai/DeepSeek-V4-Flash
200 tok/s
Qwen 3.5 35B
Qwen/Qwen3.5-35B-Instruct
150 tok/s
THROUGHPUT

40 tok/s

CONTEXT

128K

INPUT / 1M TOK

₹60

OUTPUT / 1M TOK

₹180

> Reason through a 4-step trade settlement reconciliation.
← Step 1 · match TRD-2451 → bank ref BR-9821 (amount Δ ₹0). Step 2 · flag broker fee mismatch on TRD-2452 …
OpenAI-compatible
streaming
JSON mode
FP8 quant
tool-calling
THROUGHPUT

50 tok/s

CONTEXT

200K

INPUT / 1M TOK

₹50

OUTPUT / 1M TOK

₹150

> Summarise this 80-page RFP into three bullet points.
  • Vendor must support DPDP-aligned residency
  • 99.95% SLA with 30-min credit clause
  • Decision by Q3
OpenAI-compatible
streaming
JSON mode
BF16 quant
tool-calling
vision
THROUGHPUT

70 tok/s

CONTEXT

128K

INPUT / 1M TOK

₹35

OUTPUT / 1M TOK

₹105

> Generate a JSON schema for a kirana invoice.
← { "type":"object", "properties":{ "items":[…], "gst":{ "type":"number" } } }
OpenAI-compatible
streaming
JSON mode
FP8 quant
tool-calling
THROUGHPUT

120 tok/s

CONTEXT

1M

INPUT / 1M TOK

₹28

OUTPUT / 1M TOK

₹84

> Translate this 30-min meeting transcript and tag action items.
  • Owner: Priya · Action: ship payments hot-fix by Fri
  • Owner: Rahul · Action: align with legal on DPDP scope
OpenAI-compatible
streaming
JSON mode
FP8 quant
tool-calling
vision
THROUGHPUT

200 tok/s

CONTEXT

64K

INPUT / 1M TOK

₹12

OUTPUT / 1M TOK

₹36

> Write a Python function to dedupe rows by fuzzy match.
← def dedupe_fuzzy(rows, key, threshold=0.92): # rapidfuzz.process.extract → cluster → keep canonical
OpenAI-compatible
streaming
JSON mode
INT8 quant
tool-calling
THROUGHPUT

150 tok/s

CONTEXT

128K

INPUT / 1M TOK

₹18

OUTPUT / 1M TOK

₹54

> Classify this UPI complaint by type and urgency.
← category: refund_pending · severity: medium · sla: 24h
OpenAI-compatible
streaming
JSON mode
FP8 quant
tool-calling
vision

** tok/sec on shared endpoints is subject to differ based on real-time traffic. Opt for dedicated endpoints for guaranteed performance.

Every request, instrumented.
Every layer, visible.

samaira.ai/observability

Requests / sec

30,418

4.2% vs 5m

Tokens / sec

1,415,273

6.1%

Error rate

0.05%

— stable

Cost / 1M tok

$0.66

1.8%

Request Latency

window: 5m · 1s buckets
Success rate

99.93%

Failover events

417

Retries

1703

GPU utilization

78%

Supercharge your AI agents with compliance and infinite scale.

Frontier inference, inside the boundary, pay in INR.

Run the Samaira Stack On Prem

End-to-end GPU orchestration, inside your infrastructure.

End-to-End GPU Orchestration

Full-stack GPU cluster management — provisioning, scheduling, and scaling on your own hardware.

Agentic Tuner

AI-driven auto-tuner that maximizes GPU utilization and inference performance for your workload mix.

Agentic Sandbox

Secure execution environment for multi-step agent workflows and tool-use chains on private infra.

TEE Support & Observability

Hardware-level trust with Trusted Execution Environments plus full-stack observability built in.

What's coming next.

Coming Soon

TEE Support

Confidential compute for workload isolation and hardware-level trust. Encryption in use, attestation by default.

Coming Soon

Dedicated Endpoints

Reserved capacity, custom scaling policies, and endpoint-level monitoring for predictable production workloads.

Coming Soon

Agentic Sandbox

Secure, sandboxed execution environment for agent workflows, tool use, and multi-step reasoning chains.

Enterprise AI inference,
built for India.

Secure, fast, and fully visible. Talk to us about bringing your
inference workloads inside the boundary.

$ curl https://inference.samaira.ai/openai/v1/chat/completions \
  -H "Authorization: Bearer $SAMAIRA_KEY" \
  -H "Content-Type: application/json" \
  -d '{"model":"MiniMaxAI/MiniMax-M2.7","messages":[{"role":"user","content":"Hello, India."}],"stream":false}'

{ "id": "chatcmpl-RGEzCmIB...", "object": "chat.completion", "choices": [{ "message": { "role": "assistant", "content": "Namaste! How can I help you today?" } }], "usage": { "prompt_tokens": 44, "completion_tokens": 12, "total_tokens": 56 } }
Scroll to Top