Key capabilities
- OpenAI-compatible — Works as a drop-in replacement with the OpenAI SDK
- 256K context — 262,144 tokens for large documents and multi-turn conversations
- Multimodal input — Accepts text, image, and video content
- Thinking mode — Toggle via the
thinkingparameter; returnsreasoning_contentand supports Preserved Thinking - Long-horizon coding — More reliable across languages (Rust, Go, Python) and tasks (front-end, ops, performance)
- Rich features — Tool Calls (function calling), JSON Mode, Partial Mode, web search, and automatic context caching
Quick example
Note:image_urlandvideo_urlaccept two formats: a base64 data URI (data:image/png;base64,.../data:video/mp4;base64,...) or a file reference (ms://<file_id>). Tokens served from the context cache are reported inusage.prompt_tokens_details.cached_tokens.
Parameters
API Reference
View the interactive API playground for Kimi-K2.6.