Models

16 modelos

DeepSeek V4 Flash Vision Exp is an experimental vision-enabled version of [DeepSeek V4 Flash 0731](https://openrouter.ai/deepseek/deepseek-v4-flash-0731) from DeepSeek, adding image understanding while matching the base model on text capabilities including agents,...

por deepseek21 de ago. de 20261,05 mi de contexto₩ 386 / ₩ 1.158 · 1M
1,08 tri tokens por semana

DeepSeek V4 Pro 0813 is a large-scale mixture-of-experts model from DeepSeek. This is the GA release of DeepSeek V4 Pro.

por deepseek13 de ago. de 20261,05 mi de contexto₩ 1.967 / ₩ 5.900 · 1M

DeepSeek V4 Pro 0813 is a large-scale mixture-of-experts model from DeepSeek. This is the GA release of DeepSeek V4 Pro.

por deepseek13 de ago. de 20261,05 mi de contexto₩ 2.317 / ₩ 6.950 · 1M
11,29 tri tokens por semana

DeepSeek V4 Flash 0731 is a sparse mixture-of-experts model from DeepSeek, with 13B active parameters out of 284B total. This re-post-trained revision is suited for coding, reasoning, and agent workflows....

por deepseek31 de jul. de 20261,31 mi de contexto₩ 114 / ₩ 316 · 1M

DeepSeek V4 Flash 0731 is a sparse mixture-of-experts model from DeepSeek, with 13B active parameters out of 284B total. This re-post-trained revision is suited for coding, reasoning, and agent workflows....

por deepseek31 de jul. de 20261,05 mi de contexto₩ 246 / ₩ 491 · 1M
1,49 tri tokens por semana

DeepSeek V4 Pro is a large-scale Mixture-of-Experts model from DeepSeek with 1.6T total parameters and 49B activated parameters, supporting a 1M-token context window. It is designed for advanced reasoning, coding,...

por deepseek24 de abr. de 20261,05 mi de contexto₩ 1.296 / ₩ 2.591 · 1M
5,18 tri tokens por semana

DeepSeek V4 Flash is an efficiency-optimized Mixture-of-Experts model from DeepSeek with 284B total parameters and 13B activated parameters, supporting a 1M-token context window. It is designed for fast inference and...

por deepseek24 de abr. de 20261,05 mi de contexto₩ 143 / ₩ 286 · 1M

DeepSeek-V3.2 is a large language model designed to harmonize high computational efficiency with strong reasoning and agentic tool-use performance. It introduces DeepSeek Sparse Attention (DSA), a fine-grained sparse attention mechanism...

por deepseek1 de dez. de 2025163,84 mil de contexto₩ 472 / ₩ 702 · 1M

DeepSeek-V3.2-Exp is an experimental large language model released by DeepSeek as an intermediate step between V3.1 and future architectures. It introduces DeepSeek Sparse Attention (DSA), a fine-grained sparse attention mechanism...

por deepseek29 de set. de 2025163,84 mil de contexto₩ 474 / ₩ 720 · 1M

DeepSeek-V3.1 Terminus is an update to [DeepSeek V3.1](/deepseek/deepseek-chat-v3.1) that maintains the model's original capabilities while addressing issues reported by users, including language consistency and agent capabilities, further optimizing the model's...

por deepseek22 de set. de 2025163,84 mil de contexto₩ 474 / ₩ 1.755 · 1M

DeepSeek-V3.1 is a large hybrid reasoning model (671B parameters, 37B active) that supports both thinking and non-thinking modes via prompt templates. It extends the DeepSeek-V3 base with a two-phase long-context...

por deepseek21 de ago. de 2025163,84 mil de contexto₩ 965 / ₩ 2.896 · 1M

May 28th update to the [original DeepSeek R1](/deepseek/deepseek-r1) Performance on par with [OpenAI o1](/openai/o1), but open-sourced and with fully open reasoning tokens. It's 671B parameters in size, with 37B active...

por deepseek29 de mai. de 2025163,84 mil de contexto₩ 878 / ₩ 3.773 · 1M

DeepSeek V3, a 685B-parameter, mixture-of-experts model, is the latest iteration of the flagship chat model family from the DeepSeek team. It succeeds the [DeepSeek V3](/deepseek/deepseek-chat-v3) model and performs really well...

por deepseek24 de mar. de 2025163,84 mil de contexto₩ 439 / ₩ 1.755 · 1M

DeepSeek R1 Distill Llama 70B is a distilled large language model based on [Llama-3.3-70B-Instruct](/meta-llama/llama-3.3-70b-instruct), using outputs from [DeepSeek R1](/deepseek/deepseek-r1). The model combines advanced distillation techniques to achieve high performance across...

por deepseek24 de jan. de 20258 mil de contexto₩ 1.404 / ₩ 1.404 · 1M

DeepSeek R1 is here: Performance on par with [OpenAI o1](/openai/o1), but open-sourced and with fully open reasoning tokens. It's 671B parameters in size, with 37B active in an inference pass....

por deepseek20 de jan. de 202564 mil de contexto₩ 1.229 / ₩ 4.388 · 1M

DeepSeek-V3 is the latest model from the DeepSeek team, building upon the instruction following and coding abilities of the previous versions. Pre-trained on nearly 15 trillion tokens, the reported evaluations...

por deepseek27 de dez. de 2024163,84 mil de contexto₩ 562 / ₩ 1.562 · 1M