API & Models
Available Models
Dedicated Gateway serves as a proxy to the full IG1 AI catalog.
LLM Models
| Model ID | Model Name | Capabilities | Max Context |
|---|---|---|---|
glm53-flash-chat | GLM 5.3 Flash Chat | Standard, high-velocity coding, code review, architectural decisions — with configurable reasoning depth | 1M tokens |
glm53-flash-chat-low | GLM 5.3 Flash Chat Low | Standard, high-velocity coding, code review, architectural decisions — with minimal reasoning | 1M tokens |
glm53-flash-agent | GLM 5.3 Flash Agent | Deep debugging, algorithm design, complex system modeling, agentic workflows — with configurable reasoning depth | 1M tokens |
glm53-flash-agent-low | GLM 5.3 Flash Agent Low | Deep debugging, algorithm design, complex system modeling, agentic workflows — with minimal reasoning | 1M tokens |
qwen35-122b-a10b-instruct-general | Qwen3.5-122B-A10B | Fast, accurate, and ready for any mixed workload. | 256K tokens |
qwen35-122b-a10b-instruct-creative | Qwen3.5-122B-A10B Creative | Perfect for copywriting, storytelling, and brainstorming sessions that need a spark. | 256K tokens |
qwen35-122b-a10b-thinking-general | Qwen3.5-122B-A10B Thinking | Engages deep reasoning | 256K tokens |
qwen35-122b-a10b-thinking-coding | Qwen3.5-122B-A10B Thinking Coder | pair programmer for algorithms, refactoring, and system design | 256K tokens |
qwen36-35b-a3b-instruct | Qwen 3.6 35B | Lightweight general-purpose model — text, code, and OCR/vision for fast, cost-effective tasks | 256K tokens |
qwen36-35b-a3b-thinking-general | Qwen 3.6 35B Thinking | Thinking mode activated for lightweight reasoning tasks | 256K tokens |
qwen36-35b-a3b-thinking-coding | Qwen 3.6 35B Thinking Coder | Thinking mode with coding-optimized sampling parameters | 256K tokens |
LLM models support images up to 4096 × 4096 pixels. However, processing images high-resolution images (more than 4K) can significantly increase:
- Computing time: Varies based on resolution and image complexity
- Token consumption: Higher resolutions use more input tokens
Recommendation: For optimal performance and cost efficiency, resize images to 2K maximum (2560 × 1440 pixels) before sending.
Select the right LLM model
GLM 5.3 Flash
Choose it when you're building software, running multi-agent workflows, or tackling complex engineering tasks that demand frontier-level reasoning.
GLM 5.3 Flash is a development-specialized model built for software engineering tasks — from high-velocity coding and code review to deep debugging and architectural decisions. It supports text, code, and OCR/vision (images, diagrams, tables) and offers configurable reasoning depth via the reasoning_effort parameter.
The four model variants are fixed combinations of two routing-layer settings. Users pick the variant; they cannot change clear_thinking, only reasoning_effort:
| Variant | clear_thinking | reasoning_effort | Best for |
|---|---|---|---|
glm53-flash-chat | true (thinking cleared / hidden) | low / high / max (default) | General development, chat completions |
glm53-flash-chat-low | true (thinking cleared / hidden) | Locked to low | Fast responses, code review |
glm53-flash-agent | false (thinking preserved / exposed) | low / high / max (default) | Tool-use loops, autonomous agents |
glm53-flash-agent-low | false (thinking preserved / exposed) | Locked to low | Fast agentic tasks |
On glm53-flash-chat and glm53-flash-agent, you can control reasoning depth with the reasoning_effort parameter:
| Value | Behavior |
|---|---|
low | Minimal reasoning. Fastest responses, lowest token usage. Best for straightforward queries, code review, and simple refactoring. |
high | Thorough reasoning. Deeper analysis for debugging, system design, and complex problem solving. |
max | Maximum reasoning (default). Exhaustive analysis for algorithms, architecture decisions, and multi-step planning. |
Set reasoning_effort in the request body alongside model and messages.
On the -low variants (glm53-flash-chat-low and glm53-flash-agent-low), reasoning_effort is locked to low and cannot be overridden.
Qwen 3.5 122B
Choose it when you need a reliable, high-performance model that handles everything from creative writing to visual analysis — without breaking the budget.
Qwen 3.5 122B is the Swiss Army knife of our model lineup. It does it all, and it does it cost-effectively. With four tuned variants, you get precisely the right personality for every task.
The Creative and Thinking Coder variants enforce specific sampling parameters on the backend. If you need full control over sampling, use the base instruct-general or thinking-general models instead.
Qwen 3.6 35B
Choose it when you need a fast, cost-effective model for everyday tasks — from quick answers and simple coding to lightweight creative work.
Qwen 3.6 35B is a compact 35B-parameter general-purpose model that punches above its weight. It supports text, code, and OCR/vision and handles the majority of day-to-day AI tasks at a fraction of the cost of larger models.
The three variants are tuned for different use cases:
qwen36-35b-a3b-instruct— The base model. No thinking mode. Best for straightforward chat, quick lookups, and simple generation tasks.qwen36-35b-a3b-thinking-general— Thinking mode activated. The model reasons before answering. Best for lightweight analysis, step-by-step explanations, and basic problem solving.qwen36-35b-a3b-thinking-coding— Thinking mode with coding-optimized sampling defaults. Best for quick coding assistance, snippets, and simple refactoring. If you need full control over sampling parameters, useqwen36-35b-a3b-thinking-generalinstead.
Details and usage example available under API documentation.
Use your dedicated URL Endpoint instead of api.ig1.ai
Embedding & Reranking models
| Model ID | Model Name | Dimensions | Max Context | IG1 AI Model Name |
|---|---|---|---|---|
qwen3-vl-embedding-8b | Qwen3 VL Embedding | 128 – 4096 (configurable) | 32K tokens | Qwen3-VL-Embedding 8B |
bge-m3 | BGE-M3 | 1024 (fixed) | 8K tokens | BGE-m3 |
bge-reranker-v2-m3 | BGE-Reranker-v2-M3 | N/A (reranker) | BGE-Reranker-v2-m3 |
Image models
| Model ID | Model Name | Capabilities | IG1 AI Model Name |
|---|---|---|---|
qwen-image | Qwen Image | Image generation | Qwen Image |
qwen-image-pe | Qwen Image PE | Image generation with prompt enhancement | Qwen Image PE |
qwen-image-edit | Qwen Image Edit | Image editing | Qwen Image Edit |
qwen-image-edit-pe | Qwen Image Edit PE | Image editing with prompt enhancement | Qwen Image Edit PE |
Virtual Models
Virtual Models are custom named models on top of IG1 AI models. By default IG1 teams deploy following models on each dedicated gateways.
| Virtual Model ID | IG1 AI Model ID |
|---|---|
coding-max | glm53-flash-agent-low |
coding-max-thinking | glm53-flash-agent |
general-max | glm53-flash-chat-low |
general-max-thinking | glm53-flash-chat |
coding-pro-thinking | qwen35-122b-a10b-thinking-coding |
general-pro | qwen35-122b-a10b-instruct-general |
general-pro-thinking | qwen35-122b-a10b-thinking-general |
coding-std-thinking | qwen36-35b-a3b-thinking-coding |
general-std | qwen36-35b-a3b-instruct |
general-std-thinking | qwen36-35b-a3b-thinking-general |
image-pro | qwen-image-pe |
image-pro-edit | qwen-image-edit-pe |
API Supported endpoints
Dedicated Gateway supports following API endpoints from IG1 AI API documentation:
- GET
/v1/models - POST
/v1/chat/completions - POST
/v1/embeddings - POST
/cohere/v2/embed - POST
/cohere/v2/rerank - POST
/v1/images/generations - POST
/v1/images/edits - POST
/tokenize - GET
/pricing
Please refer to api reference for all details.
Replace API reference endpoint URL from example by your own Dedicated Gateway endpoint.
Example:
Replace https://api.ig1.ai by https://api.demo.ig1.ai if you use the Demo Endpoint.
API Unsupported endpoints
Dedicated Gateway do not support following endpoints:
- POST
/document - GET
/usage - GET
/usage/history
You should use your API Master Key and direct request to https://api.ig1.ai instead if you need to call these endpoints.