LLM Models
LLM Models
| Model Name | Model ID | Capabilities | Max Context |
|---|---|---|---|
| Kimi K2.6 1T A32B Instant | kimi-k26-instant | Standard, High-velocity coding, code review, architectural decisions | 256K tokens |
| Kimi K2.6 1T A32B Thinking | kimi-k26-thinking | Deep debugging, algorithm design, complex system modeling | 256K tokens |
| Kimi K2.6 1T A32B Thinking Preserve | kimi-k26-thinking-preserve | Autonomous agents with preserved reasoning traces across turns | 256K tokens |
| Qwen3.5-122B-A10B | qwen35-122b-a10b-instruct-general | Fast, accurate, and ready for any mixed workload. | 256K tokens |
| Qwen3.5-122B-A10B Creative | qwen35-122b-a10b-instruct-creative | Perfect for copywriting, storytelling, and brainstorming sessions that need a spark. | 256K tokens |
| Qwen3.5-122B-A10B Thinking | qwen35-122b-a10b-thinking-general | Engages deep reasoning | 256K tokens |
| Qwen3.5-122B-A10B Thinking Coder | qwen35-122b-a10b-thinking-coding | pair programmer for algorithms, refactoring, and system design | 256K tokens |
LLM models support images up to 4096 × 4096 pixels. However, processing images high-resolution images (more than 4K) can significantly increase:
- Computing time: Varies based on resolution and image complexity
- Token consumption: Higher resolutions use more input tokens
Recommendation: For optimal performance and cost efficiency, resize images to 2K maximum (2560 × 1440 pixels) before sending.
Select the right model
Kimi K2.6
Choose it when you're building software, running multi-agent workflows, or tackling complex engineering tasks that demand frontier-level reasoning.
Kimi K2.6 delivers Claude Opus 4.6-class performance with a focus on agentic coding. It is perfect to architects solutions, orchestrates multi-agent collaborations, and thinks several steps ahead.
All Kimi K2.6 virtual models carry a suffix (-instant, -thinking, -thinking-preserve) to indicate that the request is altered by the IG1 AI routing layer. None of them provide direct, unmodified access to the raw model.
Kimi K2.6 is a thinking model by default. The -instant suffix alters the request to force the model to not produce thinking traces and jump to the final response directly, like a non thinking model.
Kimi K2.6 supports two distinct thinking behaviors:
Interleaved thinking (default in kimi-k26-thinking)
- The model can produce reasoning, then call a tool, then continue reasoning — all within a single conversational turn.
- If the client sends back the reasoning content during this interleaved flow, the model can continue its thinking trace using the tool results. This avoids forcing the model to restart its thinking trace from scratch, using a new tool response as a fresh starting point. This greatly enhances model coherence and reasoning quality; therefore, clients are strongly encouraged to echo back thinking traces. Thinking traces from previous conversational turns—where the assistant has already produced its final answer—are automatically dropped and excluded from input token counts.
- This is already the default behavior in the chat template and works as long as the client returns the reasoning content.
Preserved thinking (kimi-k26-thinking-preserve)
- The model preserves its full reasoning trace across multiple conversational turns.
- This prevents an autonomous agent from emitting a hypothesis, invalidating it, and then re-emitting the same hypothesis in the next turn because it no longer "remembers" it.
- Over multiple turns, this reduces redundant thinking and makes the agent more performant.
- Purpose-built for autonomous agents and long-running workflows.
Qwen 3.5 122B
Choose it when you need a reliable, high-performance model that handles everything from creative writing to visual analysis — without breaking the budget.
Qwen 3.5 122B is the Swiss Army knife of our model lineup. It does it all, and it does it cost-effectively. With four tuned variants, you get precisely the right personality for every task. :::
The Creative and Thinking Coder variants enforce specific sampling parameters on the backend. If you need full control over sampling, use the base instruct-general or thinking-general models instead.
Kimi K2.6 examples
kimi-k26-instant
Use this variant when you need fast, deterministic responses without exposed reasoning traces. Ideal for code review, linting, or straightforward refactoring.
curl https://api.ig1.ai/v1/chat/completions \
-H "Authorization: Bearer $IG1_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "kimi-k26-instant",
"messages": [
{
"role": "user",
"content": "Review this Python function for edge cases:\n\ndef divide(a, b):\n return a / b"
}
]
}'
Response (no reasoning trace produced by the model, the model directly output its final answer):
The function `divide(a, b)` has a critical edge case: division by zero.
When `b == 0`, Python raises a `ZeroDivisionError`.
Suggested fix:
def divide(a, b):
if b == 0:
raise ValueError("Cannot divide by zero")
return a / b
kimi-k26-thinking
Use this variant when a single complex problem requires reasoning interleaved with tool calls within one assistant turn. The model thinks, calls a tool, then continues thinking — all before producing final content.
Scenario: Debugging a failing test in a large codebase.
curl https://api.ig1.ai/v1/chat/completions \
-H "Authorization: Bearer $IG1_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "kimi-k26-thinking",
"messages": [
{
"role": "user",
"content": "Test `test_auth_token` is failing. Here is the traceback: [...]"
}
],
"tools": [
{
"type": "function",
"function": {
"name": "read_file",
"description": "Read a file from the repository",
"parameters": {
"type": "object",
"properties": {
"path": {"type": "string"}
}
}
}
}
]
}'
What happens inside a single turn:
- Reasoning — The model analyzes the traceback and hypothesizes the bug is in
auth.py. - Tool call — It calls
read_file(path="auth.py"). - Reasoning — With the file content, it notices
jwt.decodeis missing thealgorithmsparameter. - Final content — It returns the fix:
# Before (broken)
payload = jwt.decode(token, SECRET)
# After (fixed)
payload = jwt.decode(token, SECRET, algorithms=["HS256"])
All four steps happen within a single assistant turn — the reasoning trace is internal and the tool call is part of the same response.
kimi-k26-thinking-preserve
Use this variant when building agents that operate across multiple conversational turns and must retain their reasoning history to avoid redundant work.
Scenario: An autonomous agent investigating a production incident across several tool-use rounds and multiple conversational turns.
A single conversational turn can contain multiple API turns when tool use is involved. The distinction matters:
- API turns are the individual request/response exchanges between the client and the API (e.g., assistant thinks → calls a tool → continues thinking).
- Conversational turns are the logical back-and-forth between the user and the assistant. A new conversational turn starts when the user sends a new message after the assistant has produced its final answer.
Conversational Turn 1 — User asks the agent to investigate
{
"model": "kimi-k26-thinking-preserve",
"messages": [
{"role": "user", "content": "Service `payment-api` returned 502s at 14:03. Investigate."}
],
"tools": [...]
}
API Turn 1.1 — Assistant thinks and calls a tool (still within conversational Turn 1):
{
"role": "assistant",
"reasoning": "Hypothesis A: The 502 errors suggest the upstream service is unreachable.\nI should check the ingress logs and the pod health in the payment namespace.",
"content": "I'll investigate the 502 errors. Let me check the ingress logs and pod status."
}
User echoes back thinking and sends tool result (continuing conversational Turn 1):
{
"model": "kimi-k26-thinking-preserve",
"messages": [
{"role": "user", "content": "Service `payment-api` returned 502s at 14:03. Investigate."},
{"role": "assistant", "reasoning": "Hypothesis A: The 502 errors suggest the upstream service is unreachable.\nI should check the ingress logs and the pod health in the payment namespace.", "content": "I'll investigate the 502 errors. Let me check the ingress logs and pod status."},
{"role": "tool", "content": "Ingress logs: upstream connect error to 10.0.4.17:8080. Pod `payment-api-7d9f4b` status: CrashLoopBackOff."}
],
"tools": [...]
}
API Turn 1.2 — Assistant continues thinking and responds (still within conversational Turn 1):
{
"role": "assistant",
"reasoning": "Hypothesis A confirmed: upstream is unreachable because the pod is in CrashLoopBackOff.\nNext hypothesis B: The pod is crashing due to an OOMKill or a failed readiness probe.\nI need to check the pod events and recent logs.",
"content": "The pod `payment-api-7d9f4b` is in `CrashLoopBackOff`. I'm checking events and logs to find the root cause."
}
So far, this is identical to kimi-k26-thinking: interleaved thinking within a single conversational turn. The assistant had access to its previous thinking because the user echoed it back.
Conversational Turn 2 — User sends a follow-up (new conversational turn): The user now sends a new message. In preserve mode, ALL previous thinking traces are carried forward. The user must echo back the full conversation history, including all reasoning content.
{
"model": "kimi-k26-thinking-preserve",
"messages": [
{"role": "user", "content": "Service `payment-api` returned 502s at 14:03. Investigate."},
{"role": "assistant", "reasoning": "Hypothesis A: The 502 errors suggest the upstream service is unreachable.\nI should check the ingress logs and the pod health in the payment namespace.", "content": "I'll investigate the 502 errors. Let me check the ingress logs and pod status."},
{"role": "tool", "content": "Ingress logs: upstream connect error to 10.0.4.17:8080. Pod `payment-api-7d9f4b` status: CrashLoopBackOff."},
{"role": "assistant", "reasoning": "Hypothesis A confirmed: upstream is unreachable because the pod is in CrashLoopBackOff.\nNext hypothesis B: The pod is crashing due to an OOMKill or a failed readiness probe.\nI need to check the pod events and recent logs.", "content": "The pod `payment-api-7d9f4b` is in `CrashLoopBackOff`. I'm checking events and logs to find the root cause."},
{"role": "user", "content": "Any updates on the root cause?"}
],
"tools": [...]
}
Assistant response (conversational Turn 2):
{
"role": "assistant",
"reasoning": "From previous turn, I established:\n- Hypothesis A confirmed: upstream unreachable due to CrashLoopBackOff\n- Hypothesis B pending: check for OOMKill or failed readiness probe\n\nChecking pod events... The pod was killed with OOMKilled exit code 137.\nHypothesis B confirmed: memory limit too low.\n\nNext hypothesis C: Was there a memory spike or a leak?\nLooking at memory usage patterns...",
"content": "The root cause is an OOMKill (exit code 137). The pod exceeded its memory limit. I recommend increasing the memory limit from 512Mi to 1Gi and adding a gradual memory leak test."
}
Why this matters: With kimi-k26-thinking, the chat template would discard the thinking traces from conversational Turn 1 at the start of Turn 2. The agent would have to re-formulate "Hypothesis A: upstream is unreachable" and "Hypothesis B: check for OOMKill" from scratch, wasting tokens re-discovering what it already knew. With kimi-k26-thinking-preserve, the reasoning trace is carried forward across conversational turns, making the agent faster and cheaper across long sessions.
For implementation details on the underlying mechanism, see the Kimi K2.6 model card.