Qwen3-Coder-30B-A3B-Instruct-P-EAGLE
323
2
llama
by
amazon
Other
OTHER
30B params
New
323 downloads
Early-stage
Edge AI:
Mobile
Laptop
Server
68GB+ RAM
Mobile
Laptop
Server
Quick Summary
AI model with specialized capabilities.
Device Compatibility
Mobile
4-6GB RAM
Laptop
16GB RAM
Server
GPU
Minimum Recommended
28GB+ RAM
Code Examples
Usagetextvllm
vllm serve \
--model Qwen/Qwen3-Coder-30B-A3B-Instruct \
--tensor-parallel-size 1 \
--max-model-len 16384 \
--speculative-config '{"method": "eagle3", "model": "amazon/Qwen3-Coder-30B-A3B-Instruct-P-EAGLE", "num_speculative_tokens": 10, "parallel_drafting": true}' \
--no-enable-prefix-caching \
--async-schedulingtextvllm
| K | Acceptance Length |
|---|-------------------|
| 4 | 4.30 |
| 10 | 6.66 |
| 18 | 7.51 |
vLLM bench command is shown as below.Deploy This Model
Production-ready deployment in minutes
Together.ai
Instant API access to this model
Production-ready inference API. Start free, scale to millions.
Try Free APIReplicate
One-click model deployment
Run models in the cloud with simple API. No DevOps required.
Deploy NowDisclosure: We may earn a commission from these partners. This helps keep LLMYourWay free.