PLATFORM CAPABILITIES

Built for production AI applications.

Eight verified, specialized infrastructure capabilities covering high-context reasoning, precise technical translation, sub-200ms speech synthesis, and full-duplex voice cloning.

๐Ÿ’ฌ
Gateway API

Chat & RAG at Scale

OpenAI SDK-compatible completions with streaming, function calling, tool use, and vision inputs. Includes high-context reasoning and private on-premise local execution lanes.

Host: https://ai.ollalink.com/v1
Auth: Authorization: Bearer <key>
Models: chat, code, local
๐ŸŒ
Gateway API

Technical Translation

Machine translation designed for engineering, IT, and networking documentation. Preserves IP addresses, CIDR blocks, port numbers, and Cisco/Juniper CLI commands in Latin script.

Host: https://ai.ollalink.com/v1
Auth: Authorization: Bearer <key>
Model: translate (Indic + European)
๐ŸŽ™๏ธ
Speech GPU

Speech-to-Text (STT)

Whisper-optimized ASR pipeline offering standard audio transcription, diarization, and high-accuracy realtime streaming over low-overhead WebSockets.

Host: gpu-*.ollalink.com
Auth: X-NH-GPU-Key: <key>
Latency: ~180ms streaming chunk
๐Ÿ”Š
Speech GPU

Text-to-Speech (TTS)

50-voice catalog loudness-normalized to โˆ’23 LUFS (EBU R128 standard). Zero loudness bias, sub-200ms audio time-to-first-byte, and native Opus chunk streaming.

Host: gpu-*.ollalink.com
Auth: X-NH-GPU-Key: <key>
Voices: 50 voices (EN, HI, KN, etc.)
๐Ÿ‡ฎ๐Ÿ‡ณ
Voice Studio

Indic Voice Studio

Phonetically accurate voice models tuned across Hindi, Kannada, Tamil, Telugu, and Indian English with deterministic rules for acronym initialisms (e.g. "TCP IP").

Loudness: โˆ’23 LUFS (0.24 spread)
Languages: English, Hindi, Kannada, Tamil
Rule: Joined initialisms, no slashes
๐ŸŽญ
Voice Studio

Voice Cloning & Design

Zero-shot speaker cloning and synthetic voice persona generation from 30 seconds of clean reference audio. Deploy unique corporate vocal identities instantly.

Sample Required: 30s clean WAV/MP3
Turnaround: Instant embeddings
Security: Acoustic watermarking
โšก
Live Streaming

Realtime Live Translation

Stream spoken audio into a live duplex WebSocket and receive real-time translated audio chunks and textual captions with sub-second turnaround.

Protocol: WSS WebSocket duplex
Format: Raw PCM / Opus stream
Output: Dual audio + JSON captions
๐Ÿ‘๏ธ
Gateway API

Vision & Multimodal Parsing

High-resolution visual comprehension for enterprise schematics. Extract topology nodes, IP allocations, and structured interface tables from architectural diagrams.

Endpoint: /v1/chat/completions
Payload: Image URL or Base64 URI
Response: Validated JSON schema