Janus-Pro
Last updated 2026-09-25. Janus-Pro (Multimodal)
An advanced open multimodal foundation model that decouples visual understanding from visual generation. Utilizes a SigLIP-L visual encoder for deep scene understanding while employing discrete Vector Quantization (VQ) codebooks and parallel prediction heads for high-fidelity text-to-image synthesis.
- Parameters
- 7B / 1B Distillations; active Dense 7B Backbone
- Architecture
- Decoupled Vision: SigLIP-L Understanding + Discrete VQ Image Generation
- Context
- 32K Tokens
- KV cache
- Not listed
- Peak input / 1M (cache miss)
- 0.3
- Peak output / 1M
- 0.9
Price source: DeepSeek Models & Pricing.
Sourced benchmarks
- MMBench: 85.2% — Comprehensive multimodal benchmark assessing perception and reasoning.
- GenEval: 0.81score — Systematic evaluation of compositionality in generative image models.
- POPE (Object Hallucination): 89.4% — Polling-based object-probing evaluation benchmark.
- Seed-Bench 2: 81.6% — Multi-level visual comprehension and reasoning assessment.
All models · API pricing · FAQ