Skip to main content

LLM Models

LLM Models​

Model NameModel IDCapabilitiesMax Context
GLM 5.3 Flash Chat Lowglm53-flash-chat-lowStandard, high-velocity coding, code review, architectural decisions — with minimal reasoning1M tokens
GLM 5.3 Flash Chatglm53-flash-chatStandard, high-velocity coding, code review, architectural decisions — with configurable reasoning depth1M tokens
GLM 5.3 Flash Agent Lowglm53-flash-agent-lowDeep debugging, algorithm design, complex system modeling, agentic workflows — with minimal reasoning1M tokens
GLM 5.3 Flash Agentglm53-flash-agentDeep debugging, algorithm design, complex system modeling, agentic workflows — with configurable reasoning depth1M tokens
Qwen3.5-122B-A10Bqwen35-122b-a10b-instruct-generalFast, accurate, and ready for any mixed workload.256K tokens
Qwen3.5-122B-A10B Creativeqwen35-122b-a10b-instruct-creativePerfect for copywriting, storytelling, and brainstorming sessions that need a spark.256K tokens
Qwen3.5-122B-A10B Thinkingqwen35-122b-a10b-thinking-generalEngages deep reasoning256K tokens
Qwen3.5-122B-A10B Thinking Coderqwen35-122b-a10b-thinking-codingpair programmer for algorithms, refactoring, and system design256K tokens
Qwen 3.6 35Bqwen36-35b-a3b-instructLightweight general-purpose model — text, code, and OCR/vision for fast, cost-effective tasks256K tokens
Qwen 3.6 35B Thinkingqwen36-35b-a3b-thinking-generalThinking mode activated for lightweight reasoning tasks256K tokens
Qwen 3.6 35B Thinking Coderqwen36-35b-a3b-thinking-codingThinking mode with coding-optimized sampling parameters256K tokens
Image resolution limit

LLM models support images up to 4096 × 4096 pixels. However, processing images high-resolution images (more than 4K) can significantly increase:

  • Computing time: Varies based on resolution and image complexity
  • Token consumption: Higher resolutions use more input tokens

Recommendation: For optimal performance and cost efficiency, resize images to 2K maximum (2560 × 1440 pixels) before sending.


Select the right model​

GLM 5.3 Flash​

Choose it when you're building software, running multi-agent workflows, or tackling complex engineering tasks that demand frontier-level reasoning.

GLM 5.3 Flash is a development-specialized model built for software engineering tasks — from high-velocity coding and code review to deep debugging and architectural decisions. It supports text, code, and OCR/vision (images, diagrams, tables) and offers configurable reasoning depth via the reasoning_effort parameter.

The four model variants are fixed combinations of two routing-layer settings. Users pick the variant; they cannot change clear_thinking, only reasoning_effort:

Variantclear_thinkingreasoning_effortBest for
glm53-flash-chattrue (thinking cleared / hidden)low / high / max (default)General development, chat completions
glm53-flash-chat-lowtrue (thinking cleared / hidden)Locked to lowFast responses, code review
glm53-flash-agentfalse (thinking preserved / exposed)low / high / max (default)Tool-use loops, autonomous agents
glm53-flash-agent-lowfalse (thinking preserved / exposed)Locked to lowFast agentic tasks
Reasoning effort

On glm53-flash-chat and glm53-flash-agent, you can control reasoning depth with the reasoning_effort parameter:

ValueBehavior
lowMinimal reasoning. Fastest responses, lowest token usage. Best for straightforward queries, code review, and simple refactoring.
highThorough reasoning. Deeper analysis for debugging, system design, and complex problem solving.
maxMaximum reasoning (default). Exhaustive analysis for algorithms, architecture decisions, and multi-step planning.

Set reasoning_effort in the request body alongside model and messages:

curl https://api.ig1.ai/v1/chat/completions \
-H "Authorization: Bearer $IG1_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "glm53-flash-chat",
"reasoning_effort": "high",
"messages": [
{"role": "user", "content": "Design a distributed rate-limiter..."}
]
}'

On the -low variants (glm53-flash-chat-low and glm53-flash-agent-low), reasoning_effort is locked to low and cannot be overridden.

Qwen 3.5 122B​

Choose it when you need a reliable, high-performance model that handles everything from creative writing to visual analysis — without breaking the budget.

Qwen 3.5 122B is the Swiss Army knife of our model lineup. It does it all, and it does it cost-effectively. With four tuned variants, you get precisely the right personality for every task.

Sampling parameter presets

The Creative and Thinking Coder variants enforce specific sampling parameters on the backend. If you need full control over sampling, use the base instruct-general or thinking-general models instead.

Qwen 3.6 35B​

Choose it when you need a fast, cost-effective model for everyday tasks — from quick answers and simple coding to lightweight creative work.

Qwen 3.6 35B is a compact 35B-parameter general-purpose model that punches above its weight. It supports text, code, and OCR/vision and handles the majority of day-to-day AI tasks at a fraction of the cost of larger models.

The three variants are tuned for different use cases:

  • qwen36-35b-a3b-instruct — The base model. No thinking mode. Best for straightforward chat, quick lookups, and simple generation tasks.
  • qwen36-35b-a3b-thinking-general — Thinking mode activated. The model reasons before answering. Best for lightweight analysis, step-by-step explanations, and basic problem solving.
  • qwen36-35b-a3b-thinking-coding — Thinking mode with coding-optimized sampling defaults. Best for quick coding assistance, snippets, and simple refactoring. If you need full control over sampling parameters, use qwen36-35b-a3b-thinking-general instead.

GLM 5.3 Flash examples​

glm53-flash-chat-low​

Use this variant when you need fast, deterministic responses with minimal reasoning. Ideal for code review, linting, straightforward refactoring, or any task where speed matters more than deep analysis.

curl https://api.ig1.ai/v1/chat/completions \
-H "Authorization: Bearer $IG1_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "glm53-flash-chat-low",
"messages": [
{
"role": "user",
"content": "Review this Python function for edge cases:\n\ndef divide(a, b):\n return a / b"
}
]
}'

Response (reasoning trace is internal and minimal):

The function `divide(a, b)` has a critical edge case: division by zero.
When `b == 0`, Python raises a `ZeroDivisionError`.

Suggested fix:
def divide(a, b):
if b == 0:
raise ValueError("Cannot divide by zero")
return a / b

glm53-flash-chat​

Use this variant for general development tasks where you want to control reasoning depth. The model reasons internally and produces a clean final answer.

Scenario: Debugging a failing test in a large codebase.

curl https://api.ig1.ai/v1/chat/completions \
-H "Authorization: Bearer $IG1_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "glm53-flash-chat",
"messages": [
{
"role": "user",
"content": "Test `test_auth_token` is failing. Here is the traceback: [...]"
}
],
"tools": [
{
"type": "function",
"function": {
"name": "read_file",
"description": "Read a file from the repository",
"parameters": {
"type": "object",
"properties": {
"path": {"type": "string"}
}
}
}
}
]
}'

What happens inside a single turn:

  1. Reasoning — The model analyzes the traceback and hypothesizes the bug is in auth.py.
  2. Tool call — It calls read_file(path="auth.py").
  3. Reasoning — With the file content, it notices jwt.decode is missing the algorithms parameter.
  4. Final content — It returns the fix:
    # Before (broken)
    payload = jwt.decode(token, SECRET)
    # After (fixed)
    payload = jwt.decode(token, SECRET, algorithms=["HS256"])

All four steps happen within a single assistant turn — the reasoning trace is internal and the tool call is part of the same response.

glm53-flash-agent​

Use this variant when building agents that operate across multiple conversational turns with exposed reasoning traces. The model shows its work, making it ideal for autonomous agents, long-running workflows, and complex debugging sessions where you need to inspect or preserve the reasoning chain.

Scenario: An autonomous agent investigating a production incident across several tool-use rounds and multiple conversational turns.

A single conversational turn can contain multiple API turns when tool use is involved. The distinction matters:

  • API turns are the individual request/response exchanges between the client and the API (e.g., assistant thinks → calls a tool → continues thinking).
  • Conversational turns are the logical back-and-forth between the user and the assistant. A new conversational turn starts when the user sends a new message after the assistant has produced its final answer.

Conversational Turn 1 — User asks the agent to investigate

{
"model": "glm53-flash-agent",
"messages": [
{"role": "user", "content": "Service `payment-api` returned 502s at 14:03. Investigate."}
],
"tools": [...]
}

API Turn 1.1 — Assistant thinks and calls a tool (still within conversational Turn 1):

{
"role": "assistant",
"reasoning": "Hypothesis A: The 502 errors suggest the upstream service is unreachable.\nI should check the ingress logs and the pod health in the payment namespace.",
"content": "I'll investigate the 502 errors. Let me check the ingress logs and pod status."
}

User echoes back thinking and sends tool result (continuing conversational Turn 1):

{
"model": "glm53-flash-agent",
"messages": [
{"role": "user", "content": "Service `payment-api` returned 502s at 14:03. Investigate."},
{"role": "assistant", "reasoning": "Hypothesis A: The 502 errors suggest the upstream service is unreachable.\nI should check the ingress logs and the pod health in the payment namespace.", "content": "I'll investigate the 502 errors. Let me check the ingress logs and pod status."},
{"role": "tool", "content": "Ingress logs: upstream connect error to 10.0.4.17:8080. Pod `payment-api-7d9f4b` status: CrashLoopBackOff."}
],
"tools": [...]
}

API Turn 1.2 — Assistant continues thinking and responds (still within conversational Turn 1):

{
"role": "assistant",
"reasoning": "Hypothesis A confirmed: upstream is unreachable because the pod is in CrashLoopBackOff.\nNext hypothesis B: The pod is crashing due to an OOMKill or a failed readiness probe.\nI need to check the pod events and recent logs.",
"content": "The pod `payment-api-7d9f4b` is in `CrashLoopBackOff`. I'm checking events and logs to find the root cause."
}

Conversational Turn 2 — User sends a follow-up (new conversational turn): The user now sends a new message. In agent mode, ALL previous thinking traces are carried forward. The user must echo back the full conversation history, including all reasoning content.

{
"model": "glm53-flash-agent",
"messages": [
{"role": "user", "content": "Service `payment-api` returned 502s at 14:03. Investigate."},
{"role": "assistant", "reasoning": "Hypothesis A: The 502 errors suggest the upstream service is unreachable.\nI should check the ingress logs and the pod health in the payment namespace.", "content": "I'll investigate the 502 errors. Let me check the ingress logs and pod status."},
{"role": "tool", "content": "Ingress logs: upstream connect error to 10.0.4.17:8080. Pod `payment-api-7d9f4b` status: CrashLoopBackOff."},
{"role": "assistant", "reasoning": "Hypothesis A confirmed: upstream is unreachable because the pod is in CrashLoopBackOff.\nNext hypothesis B: The pod is crashing due to an OOMKill or a failed readiness probe.\nI need to check the pod events and recent logs.", "content": "The pod `payment-api-7d9f4b` is in `CrashLoopBackOff`. I'm checking events and logs to find the root cause."},
{"role": "user", "content": "Any updates on the root cause?"}
],
"tools": [...]
}

Assistant response (conversational Turn 2):

{
"role": "assistant",
"reasoning": "From previous turn, I established:\n- Hypothesis A confirmed: upstream unreachable due to CrashLoopBackOff\n- Hypothesis B pending: check for OOMKill or failed readiness probe\n\nChecking pod events... The pod was killed with OOMKilled exit code 137.\nHypothesis B confirmed: memory limit too low.\n\nNext hypothesis C: Was there a memory spike or a leak?\nLooking at memory usage patterns...",
"content": "The root cause is an OOMKill (exit code 137). The pod exceeded its memory limit. I recommend increasing the memory limit from 512Mi to 1Gi and adding a gradual memory leak test."
}

Why this matters: With glm53-flash-chat, the chat template would discard the thinking traces from conversational Turn 1 at the start of Turn 2. The agent would have to re-formulate "Hypothesis A: upstream is unreachable" and "Hypothesis B: check for OOMKill" from scratch, wasting tokens re-discovering what it already knew. With glm53-flash-agent, the reasoning trace is carried forward across conversational turns, making the agent faster and cheaper across long sessions.