xiaomi/mimo-v2-omni
MiMo-V2-Omni is a frontier omni-modal model that natively processes image, video, and audio inputs within a unified architecture. It combines strong multimodal perception with agentic capability - visual grounding, multi...
Input⌘ / Ctrl + Enter
Try a prompt
OutputReady
Your response will appear here
Choose an example or write a prompt, then click Run.
API details and access
Model Details
ProviderXiaomi
TypeLlm
Model IDxiaomi/mimo-v2-omni
Capabilities
InputTextImage
OutputText
Context256,000
Max Output65,536
VisionSupported
Function CallingSupported
Access details
Chat CompletionsBase URLhttps://api.sandbase.ai
API Endpoint/v1/chat/completions
{
"model": "xiaomi/mimo-v2-omni",
"messages": [
{
"role": "user",
"content": "Hello"
}
],
"max_tokens": 512
}Pricing
Input$0.40 / 1M tokens
Output$2.00 / 1M tokens
Cache read$0.08 / 1M tokens

