Gemma 4 E4B
Chat and vision
gemma-4-e4bOne inline PNG, up to 1024×1024; output 640–2048 tokens including reasoning. No tools.
The API, router, PostgreSQL and queue worker run together on AWS in Cape Town. The API uses an OpenAI-compatible subset. Access is restricted to invited workspaces while we prepare operating support.
Chat and vision
gemma-4-e4bOne inline PNG, up to 1024×1024; output 640–2048 tokens including reasoning. No tools.
Semantic search
qwen3-embedding-0.6b1024 dimensions; 16 texts per batch; 2048 tokens per text.
Search relevance
qwen3-reranker-0.6b16 documents; token count includes repeated query and prompt for every document.
Speech to text
whisper-large-v3Mono PCM16 WAV, 16 kHz, up to 60 seconds; rounded up to whole seconds. African-accent quality needs correction.
Image generation
sd15One 512×512 PNG, 25 steps; queued preview with content filtering. No SDXL or video.
Text to speech
kokoro-82mEnglish only; af_heart / bf_emma; 1000 Unicode code points; WAV only, up to 60 seconds.
Install the OpenAI Node SDK and use your invitation’s key and base URL. This example runs against an enabled pilot account; the endpoint is https://api.embiro.ai/v1 and requires an invited workspace’s API key. Keys are available only after sign-in.
import OpenAI from "openai";
const client = new OpenAI({
apiKey: process.env.EMBIRO_API_KEY,
baseURL: process.env.EMBIRO_BASE_URL, // invitation URL including /v1
maxRetries: 0, // avoid replaying synchronous paid work
});
const response = await client.chat.completions.create({
model: "gemma-4-e4b",
messages: [{ role: "user", content: "Explain semantic search." }],
max_tokens: 1024, // includes the model's reasoning budget
});
console.log(response.choices[0].message.content);Supported: model discovery, Chat Completions with streaming, inline PNG vision, embeddings, WAV transcription uploads, image generations and WAV speech. Reranking, jobs, pricing and usage are Embiro extensions. Responses API, Assistants, fine-tuning, Gemma tool calls and arbitrary file formats are not supported.
All 28 mixed requests completed with four concurrent clients, alongside priority, queue durability and cancellation recovery checks. Small-fixture medians were 1.1 seconds for chat, 6.6 seconds for vision and 2.1 seconds for speech. Measurements include scheduling wait, exclude internet latency and are not production performance guarantees.
This is one host, with scheduled sessions and operator recovery after uncertain cancellation. Whisper needs human review for African accents. Kokoro is English only. Images are a 512px preview; SDXL and video are excluded. No high availability or cross-provider standby is active.