Skip to main content

API & Models

Available Models​

Dedicated Gateway serves as a proxy to the full IG1 AI catalog.

LLM Models​

Model IDModel NameCapabilitiesMax Context
glm53-flash-chatGLM 5.3 Flash ChatStandard, high-velocity coding, code review, architectural decisions — with configurable reasoning depth1M tokens
glm53-flash-chat-lowGLM 5.3 Flash Chat LowStandard, high-velocity coding, code review, architectural decisions — with minimal reasoning1M tokens
glm53-flash-agentGLM 5.3 Flash AgentDeep debugging, algorithm design, complex system modeling, agentic workflows — with configurable reasoning depth1M tokens
glm53-flash-agent-lowGLM 5.3 Flash Agent LowDeep debugging, algorithm design, complex system modeling, agentic workflows — with minimal reasoning1M tokens
qwen35-122b-a10b-instruct-generalQwen3.5-122B-A10BFast, accurate, and ready for any mixed workload.256K tokens
qwen35-122b-a10b-instruct-creativeQwen3.5-122B-A10B CreativePerfect for copywriting, storytelling, and brainstorming sessions that need a spark.256K tokens
qwen35-122b-a10b-thinking-generalQwen3.5-122B-A10B ThinkingEngages deep reasoning256K tokens
qwen35-122b-a10b-thinking-codingQwen3.5-122B-A10B Thinking Coderpair programmer for algorithms, refactoring, and system design256K tokens
qwen36-35b-a3b-instructQwen 3.6 35BLightweight general-purpose model — text, code, and OCR/vision for fast, cost-effective tasks256K tokens
qwen36-35b-a3b-thinking-generalQwen 3.6 35B ThinkingThinking mode activated for lightweight reasoning tasks256K tokens
qwen36-35b-a3b-thinking-codingQwen 3.6 35B Thinking CoderThinking mode with coding-optimized sampling parameters256K tokens
Image resolution limit

LLM models support images up to 4096 × 4096 pixels. However, processing images high-resolution images (more than 4K) can significantly increase:

  • Computing time: Varies based on resolution and image complexity
  • Token consumption: Higher resolutions use more input tokens

Recommendation: For optimal performance and cost efficiency, resize images to 2K maximum (2560 × 1440 pixels) before sending.

Select the right LLM model​

GLM 5.3 Flash​

Choose it when you're building software, running multi-agent workflows, or tackling complex engineering tasks that demand frontier-level reasoning.

GLM 5.3 Flash is a development-specialized model built for software engineering tasks — from high-velocity coding and code review to deep debugging and architectural decisions. It supports text, code, and OCR/vision (images, diagrams, tables) and offers configurable reasoning depth via the reasoning_effort parameter.

The four model variants are fixed combinations of two routing-layer settings. Users pick the variant; they cannot change clear_thinking, only reasoning_effort:

Variantclear_thinkingreasoning_effortBest for
glm53-flash-chattrue (thinking cleared / hidden)low / high / max (default)General development, chat completions
glm53-flash-chat-lowtrue (thinking cleared / hidden)Locked to lowFast responses, code review
glm53-flash-agentfalse (thinking preserved / exposed)low / high / max (default)Tool-use loops, autonomous agents
glm53-flash-agent-lowfalse (thinking preserved / exposed)Locked to lowFast agentic tasks
Reasoning effort

On glm53-flash-chat and glm53-flash-agent, you can control reasoning depth with the reasoning_effort parameter:

ValueBehavior
lowMinimal reasoning. Fastest responses, lowest token usage. Best for straightforward queries, code review, and simple refactoring.
highThorough reasoning. Deeper analysis for debugging, system design, and complex problem solving.
maxMaximum reasoning (default). Exhaustive analysis for algorithms, architecture decisions, and multi-step planning.

Set reasoning_effort in the request body alongside model and messages.

On the -low variants (glm53-flash-chat-low and glm53-flash-agent-low), reasoning_effort is locked to low and cannot be overridden.

Qwen 3.5 122B​

Choose it when you need a reliable, high-performance model that handles everything from creative writing to visual analysis — without breaking the budget.

Qwen 3.5 122B is the Swiss Army knife of our model lineup. It does it all, and it does it cost-effectively. With four tuned variants, you get precisely the right personality for every task.

Sampling parameter presets

The Creative and Thinking Coder variants enforce specific sampling parameters on the backend. If you need full control over sampling, use the base instruct-general or thinking-general models instead.

Qwen 3.6 35B​

Choose it when you need a fast, cost-effective model for everyday tasks — from quick answers and simple coding to lightweight creative work.

Qwen 3.6 35B is a compact 35B-parameter general-purpose model that punches above its weight. It supports text, code, and OCR/vision and handles the majority of day-to-day AI tasks at a fraction of the cost of larger models.

The three variants are tuned for different use cases:

  • qwen36-35b-a3b-instruct — The base model. No thinking mode. Best for straightforward chat, quick lookups, and simple generation tasks.
  • qwen36-35b-a3b-thinking-general — Thinking mode activated. The model reasons before answering. Best for lightweight analysis, step-by-step explanations, and basic problem solving.
  • qwen36-35b-a3b-thinking-coding — Thinking mode with coding-optimized sampling defaults. Best for quick coding assistance, snippets, and simple refactoring. If you need full control over sampling parameters, use qwen36-35b-a3b-thinking-general instead.

Details and usage example available under API documentation. Use your dedicated URL Endpoint instead of api.ig1.ai

Embedding & Reranking models​

Model IDModel NameDimensionsMax ContextIG1 AI Model Name
qwen3-vl-embedding-8bQwen3 VL Embedding128 – 4096 (configurable)32K tokensQwen3-VL-Embedding 8B
bge-m3BGE-M31024 (fixed)8K tokensBGE-m3
bge-reranker-v2-m3BGE-Reranker-v2-M3N/A (reranker)BGE-Reranker-v2-m3

Image models​

Model IDModel NameCapabilitiesIG1 AI Model Name
qwen-imageQwen ImageImage generationQwen Image
qwen-image-peQwen Image PEImage generation with prompt enhancementQwen Image PE
qwen-image-editQwen Image EditImage editingQwen Image Edit
qwen-image-edit-peQwen Image Edit PEImage editing with prompt enhancementQwen Image Edit PE

Virtual Models​

Virtual Models are custom named models on top of IG1 AI models. By default IG1 teams deploy following models on each dedicated gateways.

Virtual Model IDIG1 AI Model ID
coding-maxglm53-flash-agent-low
coding-max-thinkingglm53-flash-agent
general-maxglm53-flash-chat-low
general-max-thinkingglm53-flash-chat
coding-pro-thinkingqwen35-122b-a10b-thinking-coding
general-proqwen35-122b-a10b-instruct-general
general-pro-thinkingqwen35-122b-a10b-thinking-general
coding-std-thinkingqwen36-35b-a3b-thinking-coding
general-stdqwen36-35b-a3b-instruct
general-std-thinkingqwen36-35b-a3b-thinking-general
image-proqwen-image-pe
image-pro-editqwen-image-edit-pe

API Supported endpoints​

Dedicated Gateway supports following API endpoints from IG1 AI API documentation:

  • GET /v1/models
  • POST /v1/chat/completions
  • POST /v1/embeddings
  • POST /cohere/v2/embed
  • POST /cohere/v2/rerank
  • POST /v1/images/generations
  • POST /v1/images/edits
  • POST /tokenize
  • GET /pricing

Please refer to api reference for all details.

API endpoint

Replace API reference endpoint URL from example by your own Dedicated Gateway endpoint.

Example: Replace https://api.ig1.ai by https://api.demo.ig1.ai if you use the Demo Endpoint.


API Unsupported endpoints​

warning

Dedicated Gateway do not support following endpoints:

  • POST /document
  • GET /usage
  • GET /usage/history

You should use your API Master Key and direct request to https://api.ig1.ai instead if you need to call these endpoints.