Qwen3-VL-8B-Instruct is a multimodal vision-language model from the Qwen3-VL series, built for high-fidelity understanding and reasoning across text, images, and video. It features improved multimodal fusion with Interleaved-MRoPE for long-horizon...
모달리티
이미지텍스트
입력 / 출력 단가
₩205 / ₩799
1M
컨텍스트
26.21만
출시
2025. 10. 15.
이 모델의 벤치마크 데이터가 없습니다.