Verified private deployment · invitations not open yet

Six models.
One pilot API.

The API, router, PostgreSQL and queue worker run together on AWS in Cape Town. The API uses an OpenAI-compatible subset. Access is restricted to invited workspaces while we prepare operating support.

core

Gemma 4 E4B

Chat and vision

gemma-4-e4b

One inline PNG, up to 1024×1024; output 640–2048 tokens including reasoning. No tools.

core

Qwen3 Embedding 0.6B

Semantic search

qwen3-embedding-0.6b

1024 dimensions; 16 texts per batch; 2048 tokens per text.

core

Qwen3 Reranker 0.6B

Search relevance

qwen3-reranker-0.6b

16 documents; token count includes repeated query and prompt for every document.

beta

Whisper large-v3

Speech to text

whisper-large-v3

Mono PCM16 WAV, 16 kHz, up to 60 seconds; rounded up to whole seconds. African-accent quality needs correction.

preview

Stable Diffusion 1.5

Image generation

sd15

One 512×512 PNG, 25 steps; queued preview with content filtering. No SDXL or video.

beta

Kokoro 82M

Text to speech

kokoro-82m

English only; af_heart / bf_emma; 1000 Unicode code points; WAV only, up to 60 seconds.

Use your existing SDK.

Install the OpenAI Node SDK and use your invitation’s key and base URL. This example runs against an enabled pilot account; the endpoint is https://api.embiro.ai/v1 and requires an invited workspace’s API key. Keys are available only after sign-in.

import OpenAI from "openai";

const client = new OpenAI({
  apiKey: process.env.EMBIRO_API_KEY,
  baseURL: process.env.EMBIRO_BASE_URL, // invitation URL including /v1
  maxRetries: 0, // avoid replaying synchronous paid work
});

const response = await client.chat.completions.create({
  model: "gemma-4-e4b",
  messages: [{ role: "user", content: "Explain semantic search." }],
  max_tokens: 1024, // includes the model's reasoning budget
});
console.log(response.choices[0].message.content);

Supported: model discovery, Chat Completions with streaming, inline PNG vision, embeddings, WAV transcription uploads, image generations and WAV speech. Reranking, jobs, pricing and usage are Embiro extensions. Responses API, Assistants, fine-tuning, Gemma tool calls and arbitrary file formats are not supported.

Measured together.

All 28 mixed requests completed with four concurrent clients, alongside priority, queue durability and cancellation recovery checks. Small-fixture medians were 1.1 seconds for chat, 6.6 seconds for vision and 2.1 seconds for speech. Measurements include scheduling wait, exclude internet latency and are not production performance guarantees.

This is one host, with scheduled sessions and operator recovery after uncertain cancellation. Whisper needs human review for African accents. Kokoro is English only. Images are a 512px preview; SDXL and video are excluded. No high availability or cross-provider standby is active.