Qwen3 30B A3B
qwen3-30b-a3bQwen3 is the latest generation in the Qwen large language model series, featuring both dense and Mixture-of-Experts (MoE) architectures designed to excel in reasoning, multilingual understanding, and advanced agentic tasks. A defining capability of Qwen3 is its ability to seamlessly switch between a thinking mode for complex, multi-step reasoning and a non-thinking mode for efficient, high-quality dialogue—delivering strong versatility across use cases. Compared with earlier models such as QwQ and Qwen2.5, Qwen3 demonstrates substantial performance gains in mathematics, coding, commonsense reasoning, creative writing, and interactive conversation. The Qwen3-30B-A3B variant comprises 30.5B total parameters with 3.3B activated, 48 layers, and 128 experts (with 8 activated per task). With support for up to 131K token context lengths via YaRN, it sets a new benchmark for open-source MoE models in both capability and efficiency.
- Context
- 41.0K tokens
- Endpoint
Service Status
Status information temporarily unavailable
Apertis cannot confirm the current service state. This is not a report that the model is down.
Pricing
Quick Start
Select an endpoint and copy a working example for this model.
from openai import OpenAI client = OpenAI( api_key="YOUR_API_KEY", base_url="https://api.apertis.ai/v1") response = client.chat.completions.create( model="qwen3-30b-a3b", messages=[ {"role": "user", "content": "Hello!"} ], max_tokens=1024, temperature=0.7) print(response.choices[0].message.content) # Optional: Enable context compression to reduce token usage# response = client.chat.completions.create(# model="qwen3-30b-a3b",# messages=[{"role": "user", "content": "Hello!"}],# extra_body={"compression": {"enabled": True, "model": "gpt-4.1-mini"}}# )Supported Parameters
API docsmodelmessagesmax_tokenstemperaturetop_pstreamtoolsreasoning_effortstream_optionsthinkingextra_bodyCursor IDE Model IDs
Use these namespaced identifiers in Cursor IDE to avoid conflicts with built-in models.
Compare with Other Models
See how this model compares to others from the same provider.
Qwen Plus 2025-07-28 Thinking
Qwen Plus 0728 is a hybrid reasoning model built on the Qwen3 foundation, featuring a 1M-token context window and a balanced trade-off between performance, speed, and cost.
- Context
- 1M
- Input
- $0.40/M
- Output
- $4.00/M
Tongyi DeepResearch 30B A3B
Tongyi DeepResearch is a 30B-parameter agentic model (3B active per token) built for long-horizon, deep research and information-seeking tasks. It achieves state-of-the-art results on major agentic search and reasoning benchmarks, outperforming prior models in complex multi-step problem solving. Trained with a fully automated synthetic data pipeline and advanced on-policy RL, it supports ReAct workflows and a high-performance “Heavy” mode for test-time scaling, making it well suited for advanced research agents, tool use, and intensive inference workloads.
- Context
- 131.1K
- Input
- $0.135/M
- Output
- $0.675/M
Qwen3 VL 235B A22B Thinking
Qwen3-VL-235B-A22B Thinking is a powerful multimodal model that combines advanced text generation with strong image and video understanding, optimized for STEM and math reasoning. It offers robust perception, spatial grounding, and long-form visual comprehension, and supports agent-style interactions such as multi-image dialogue, video timeline alignment, GUI control, and visual-to-code workflows. With competitive benchmark results and strong text-only ability, it’s suited for production uses like document AI, OCR, UI assistance, spatial tasks, and vision-language research.
- Context
- 131.1K
- Input
- $0.30/M
- Output
- $3.00/M
Qwen3 VL 235B A22B Instruct
Qwen3-VL-235B-A22B Instruct is an open-weight multimodal model that combines strong language generation with image and video understanding, aimed at general vision-language tasks like VQA, document parsing, chart/table extraction, and multilingual OCR. It features robust perception, spatial grounding, and long-context visual comprehension, and supports agent-style workflows such as multi-image dialogue, video timeline alignment, GUI control, and visual-to-code assistance. With competitive benchmark performance and strong text-only ability, it's well suited for production uses across document AI, OCR, UI assistance, spatial reasoning, and vision-language research.
- Context
- 131.1K
- Input
- $0.30/M
- Output
- $1.50/M