Home
Blog
AI
Article

The Real Cost of Training Frontier AI Models: What DeepSeek Revealed and What It Means for the Industry

DeepSeek-V3's reported $5.6 million training cost reshaped the conversation about frontier AI, but it covers only the final run, not the research and failed experiments behind it. This piece breaks down what AI development really costs, how MoE, MLA, FP8 training, and hardware-software co-design drove DeepSeek's efficiency, where open models still trail closed ones, and why the distillation controversy complicates the story.

The Real Cost of Training Frontier AI Models: What DeepSeek Revealed and What It Means for the Industry
Propagated AI Team
September 29, 2026
 • 
Updated 
September 29, 2026
13
 min read
The Real Cost of Training Frontier AI Models: What DeepSeek Revealed and What It Means for the Industry

What Is DeepSeek and How Is It Different from Other AI Models?

DeepSeek is a relatively new AI model that deserves attention for its cost, strong language capabilities, and focus on efficiency.

Unlike many other AI models that compete by increasing model size and computational capacity, DeepSeek takes a different approach, focusing on optimising how AI systems are trained and operated.

DeepSeek’s technical report states that training DeepSeek-V3 required approximately 2.788 million H800 GPU-hours, resulting in around $5.6 million in GPU rental costs for the official training process.

While this money is, in fact, a cost only of the final training phase, the difference is still significant, showing that advanced large language models (LLM) can be developed with far fewer resources, opening new opportunities for smaller organisations to participate in LLM development.

Apart from cost reduction, we are talking about emphasising engineering expertise. The model shows that improvements in architecture, training methods, and resource efficiency can yield highly competitive performance without depending entirely on large-scale computing. These advances could influence how companies approach LLM development and long-term maintenance by reducing some of the financial and technical barriers associated with building and operating advanced LLM.

But even more importantly, DeepSeek contributed to the rapid growth of the open AI ecosystem.

​

Key Takeaways:

  • Advanced LLM could be developed with significantly fewer resources than previously expected.
  • Transparency in technical reporting has encouraged trust, discussion, and further research among industry professionals and AI researchers.
  • DeepSeek showed the potential of reinforcement learning to boost model reasoning, encouraging greater interest in reasoning-focused training approaches.
  • Innovations in Mixture-of-Experts (MoE) architecture, Multi-head Latent Attention (MLA), FP8 mixed-precision training, and hardware-aware optimisation helped create a more efficient LLM that prioritises smarter resource utilisation over simply increasing scale.
  • DeepSeek achieves its results through several key innovations: prioritising AI learning over model size, applying a hardware–software co-design approach, improving algorithmic efficiency, and boosting model accuracy.
  • Although the future direction of the AI industry is yet unknown, DeepSeek has already influenced how organisations think about developing and scaling frontier LLM: engineering excellence may become one of the defining competitive advantages in the next generation of LLM development.

DeepSeek’s Story

At the beginning of its journey, DeepSeek AI began as a side project by Liang Wenfeng, founder of the hedge fund High-Flyer, to predict financial market trends.

Upon the release of the new DeepSeek V3 model, the creators revealed a shocking training cost of just $5.6 million, down from hundreds of millions of dollars previously required.

Bar chart comparing the estimated input and output token costs of Grok, ChatGPT-o1 Mini, Gemini 1.5 Pro, Nova Pro, DeepSeek-R1, and Llama 3.1 Nemotron 70B. DeepSeek-R1 has one of the lowest prices, while Grok and ChatGPT-o1 Mini are the most expensive.
Figure 1: Estimated input and output token pricing for leading LLM, highlighting the lower cost of DeepSeek-R1.

A graph from Statista shows the drastic change in input/output token prices among various LLM. In this context, a token refers to a small unit of text processed by a large language model. The graph compares the cost of processing input and output tokens across several models, highlighting DeepSeek’s significantly lower inference costs.

But what is important here is that this cost does not include the substantial work that made that final run possible.

Many prototype models are trained and discarded for different reasons. While evaluating different architectures, improvements often come from failed experiments rather than successful ones and every improvement is tested repeatedly before becoming part of the final model - a process that also requires time and money.

Let’s draw an analogy: confusing the cost of that final run with the cost of building the entire model is like claiming that constructing a skyscraper costs only as much as its spectacular top layer.

Illustration showing the visible training cost and the underlying stages required to develop frontier AI models.
Figure 2: Hidden foundations of LLM development.

An important factor often overlooked is the research and engineering effort required to develop, optimise, and improve the efficiency of large language models. Building an efficient LLM involves much more than using large amounts of data. It requires designing the model architecture, developing training techniques, and optimising the model's use of computing resources to process and generate human-like text.

The development of a large language model could typically involve the following events in chronological order:

  1. A lot of small-scale experiments.
  2. Architecture comparisons.
  3. Hyperparameter searches.
  4. Data filtering.
  5. Reinforcement learning research.
  6. Safety analyses.
  7. Scaling studies.
  8. Infrastructure optimization.
  9. Intermediate models.
  10. Final production training.

What is presented to the public is a final product made by dedicated researchers, rather than the gradual development process behind it. Typically, the work done to create that final product is not shown to the public, conveniently hidden to make the headlines catchier.

Flowchart illustrating the AI development process, from research and prototyping through architecture testing, data preparation, reinforcement learning, infrastructure optimization, final training, and deployment.
Figure 3: Overview of the key stages involved in developing and deploying a frontier AI model.

What Drives the High Costs of LLM Development

While efficiency eliminated a large chunk of the cost of developing an AI model, this important aspect remained largely undiscussed. We’ll explore the efficiency aspect in more detail below.

DeepSeek did not eliminate the cost of maintenance, only the training. Even though the cost of training has decreased significantly, many other factors additionally contribute to the cost of maintaining a large language model.

Very large GPU clusters can also consume enormous amounts of electricity, ending up costing tens of millions of dollars annually. Thanks to DeepSeek, some of this cost could be managed better, but the problem remains.

Donut chart showing the estimated distribution of AI development costs. The largest shares are model development and training (25%) and development team (25%), followed by data acquisition and infrastructure (15% each), testing and validation (10%), and project management and regulatory compliance (5% each).
Figure 4: Estimated percentage breakdown of costs across key stages of LLM development.

This image from Coherent Solutions shows the approximate cost ranges for different aspects of AI maintenance. This percentage can vary by model and company; however, it is an estimated cost distribution graph.

Finally, what makes DeepSeek stand out is that they released a detailed technical report that includes not only the widely discussed training cost but also the rationale behind many of their engineering decisions. This transparency is not common in the field; it shows that the people behind DeepSeek are truly proud of their work. This transparency can lead more companies to adopt it as a standard and be more truthful in the future, accelerating progress and encouraging further innovation across the field.

The Efficiency Behind the Performance

DeepSeek has a strong concentration on efficiency and prioritises smarter design over resource expenditure. We can see great LLM effectiveness as a result.

In 2024-2025, DeepSeek held leading positions in coding and quantitative reasoning, and had strong performance in reasoning and knowledge. DeepSeek had a great start and, for a while, was the highest-performing large language model.

So why was DeepSeek able to accomplish what its predecessors could not?

Rather than just being less costly, DeepSeek demonstrated the possibility of more efficient LLM training for the future of large language models.

Line chart showing DeepSeek-R1-Zero accuracy improving steadily during training. Both evaluation settings increase over time, with the red line consistently outperforming the blue line and approaching benchmark reference levels.
Figure 5: Accuracy of DeepSeek-R1-Zero during training under two evaluation settings.

This graph illustrates reinforcement training and its possible benefits for model learning. The DeepSeek R-1 model showed signs of spending more time on difficult questions, reconsidering earlier steps, and fixing mistakes before giving its final answer. This graph also demonstrates how giving the model multiple attempts can improve its performance over time. Even though this graph shows highly successful performance, it is important not to forget that the American Invitational Mathematics Examination (AIME) is only one benchmark of a large language model's overall capability.

However, by 2026, the picture had changed. The frontier had moved on to newer generations of models from OpenAI, Anthropic, and Google, making the original o1–R1 comparison outdated. One of the most discussed models was Anthropic’s Claude 3.7 Sonnet, released in February 2025. Anthropic described it as its first hybrid reasoning model and reported strong results in coding, agentic tasks and reasoning. Artificial Analysis' current Intelligence Index ranks newer models above the older o1/o3 generation. Its current rankings include Claude Opus (4.7/4.8), GPT-5.5, and other 2026 frontier models, depending on the specific model configuration and evaluation version. Thus, competition shifts again, and this time toward real-world reasoning, coding, agents and long-context work.

The Role of Engineering in LLM Efficiency

Before DeepSeek V3, many discussions about AI were focused on the model size.

It left companies working on:

  • larger parameter numbers
  • bigger datasets
  • more GPUs
  • larger data centres

However, DeepSeek challenged that assumption by introducing a new competitive advantage: engineering excellence.

Engineering innovations emerged as a more sustainable way to improve performance without a proportional increase in costs. This shift is changing what it means to be competitive in AI research.

DeepSeek uses MoE technology, which stands for Mixture-of-Experts. This means that instead of using more energy to analyse every prompt, DeepSeek activates only the experts required for a given prompt, making the model more efficient and lowering resource usage.

Accomplishments in MoE architecture, Multi-head Latent Attention (MLA), FP8 mixed-precision training, and hardware-aware optimisation all help reduce training time and cost.

On their own, these elements are important for building a successful AI model, but together they create something we have never seen before. A hyper-efficient large language model that works smarter rather than harder.

Highly successful engineering can be far more beneficial than just improving software, and DeepSeek is a prime example.

Line chart showing the average response length of DeepSeek-R1-Zero increasing steadily during training, indicating progressively longer generated responses as training advances.
Figure 6: Average response length generated by DeepSeek-R1-Zero throughout the training process.

This graph, from PromptHub, shows a gradual increase in answer length across training steps. The dark blue line shows the average response length, and the shaded region shows the variation in responses. With more steps, the model is demonstrating longer patterns of reasoning. The model was trained by reinforcement learning, and the reward system was based on producing correct answers rather than simply making them longer. However, the model realised that by taking longer to calculate the answers, it was improving their quality.

How DeepSeek Is Influencing the Future of LLM Development

There are a few innovative ways DeepSeek uses to achieve its great results:

  1. Prioritising learning over model size.
  2. Hardware–software co-design.
  3. Using algorithmic effectiveness.
  4. Accuracy upgrades.

Let’s take a closer look.

Learning & MoE

The success of V3 model would not have been possible without the DeepSeek-R1 reasoning that helped researchers further improve V3’s post-training performance. Through intensive learning, it became clear that the key to improving the model's efficiency lay in mathematical reasoning rather than in model size.

It was initially believed that improving an AI model meant adding more of everything: more data, more parameters, and more computing power. These things still matter, but DeepSeek shows that bigger isn’t always better. Instead of simply scaling up, researchers can focus on making models more efficient through better training and architecture. DeepSeek’s use of Mixture-of-Experts (MoE), for example, allows the model to use only a subset of its parameters for each task, reducing the required computation. The result is a different approach to improving LLM: rather than just making models bigger, make better use of the resources you already have.

Hardware–Software Co-Design

DeepSeek also partakes in hardware–software co-design. This means that instead of designing the models separately from the hardware, the companies can improve both. The DeepSeek-V3 model was designed around H800 NVIDIA GPUs, taking their capabilities into account. This, in turn, demonstrates how architecture and hardware can work together.

As LLM systems persist in evolving, hardware–software co-design is likely to become increasingly important, allowing future models to make more effective use of available computing resources.

The Age of Algorithmic Efficiency

As we mentioned before, the traditional assumption was that the more GPUs were used, the better the AI would be, but things are slowly starting to change. Good performance now depends on algorithms, architecture, training methods, inference optimisation, and hardware design. Modern computing does not rely only on faster processors, but prioritises software optimisation.

As we learned above, DeepSeek did not prove that hardware no longer matters. Instead, it showed that engineering innovations can dramatically improve the effectiveness of hardware use. As a result, algorithmic efficiency is becoming one of the defining measures of progress in modern LLM.

Accuracy Improvements

Another way DeepSeek changed the industry was by pushing modern reasoning models to improve. For example, instead of generating a single answer, it may generate multiple answers and select the best one to improve the model's response accuracy. DeepSeek's published evaluations present marked advances in reasoning performance, pointing to the growing importance of inference-time reasoning.

Market Shift

DeepSeek opened the possibility for smaller AI developers to compete with the current LLM leaders by making it more accessible. Improved engineering and pre-existing open-weight releases allow new researchers to build on what we already know rather than start from scratch. But even though LLM development seems more accessible, there are still factors to consider when building a large language model. These companies need researchers, infrastructure, datasets, computing expertise, financial backing, and most importantly, the ability to take risks. It might have become more competitive, but the expenses remain the same.

However, DeepSeek has changed expectations and begun setting standards companies must follow. They can no longer assume that a higher price equals a better product; they now have to put in the effort to research. Users and investors started asking vastly different questions, such as “How efficient is the model?” or “What makes it stand out from the rest?”, rather than simply “How many GPUs does it have?” These discussions drive companies to spend more time ensuring that the product represents their best work and has consumers’ needs in mind.

Open vs Closed Models: Where the Differences Remain

Despite the rapid progress of open-weight models, a gap between them and the leading closed-frontier models remains. However, this gap has never been constant. It has varied considerably across different periods, models, and benchmarks. Recent estimates place the average gap at roughly four months, while other evaluations suggest that it can be closer to eight months for longer-horizon agentic tasks.

DeepSeek-R1 was an important turning point in this development. Released in early 2025, it demonstrated reasoning performance comparable to OpenAI's o1-1217 while using a different approach centred heavily on reinforcement learning.

But soon after, the open-model ecosystem changed quite a bit. Qwen3 has combined MoE architectures with reasoning and efficient parameter activation, and models such as Kimi K2 and GLM-4.5 have pushed further into coding and agentic capabilities.

This means that newer models do not simply replicate DeepSeek's approach. Instead, they demonstrate how the wider open-model ecosystem is experimenting with many of the same principles, including reinforcement learning, MoE architectures, distillation, inference-time scaling, and hardware-efficient design, while developing their own variations.

As a result, the gap between open and closed models may be better understood as a moving target rather than a fixed six-month delay. The leading closed models continue to advance, but open-model developers are increasingly able to reproduce or approach frontier-level capabilities in much shorter periods.

Another important development is the growing use of model distillation. Distillation allows a smaller model to learn from the outputs of a more capable model, potentially transferring useful reasoning and problem-solving capabilities without requiring the same level of computing resources. Although distillation is a legitimate and widely used training technique, its use has become controversial when models are trained on the outputs of competing frontier systems.

OpenAI has accused DeepSeek of using distillation to obtain outputs from its models and other frontier models. At the same time, Anthropic has also accused DeepSeek, Moonshot, and MiniMax of conducting large-scale efforts to extract capabilities from Claude. Anthropic reported that the three companies generated more than 16 million exchanges with Claude through approximately 24,000 accounts. These allegations have increased the debate over where legitimate research and model improvement end and unauthorised replication begins.

There is also a risk that these methods will encounter new challenges related to data quality, reasoning capabilities, or inference efficiency as they scale to even larger frontier models. Continued engineering innovation will undoubtedly help address many of the challenges listed above, but it is unlikely to eliminate them.

Nevertheless, DeepSeek's success encourages competitors to invest more heavily in efficiency-focused research rather than simply increasing computational assets. Future competition will probably be defined not only by who can build the largest models, but also by who can design the smartest and most efficient ones. At the same time, the controversy surrounding distillation shows that efficiency alone may not determine AI's future. Questions about how models acquire their capabilities, whether those methods are legally and ethically acceptable, and how much original innovation is involved will become increasingly important.

The real shift may be toward achieving more with less: better algorithms, smarter training, efficient inference, specialised hardware, and more effective use of existing AI capabilities.

Conclusion

The progress we see with DeepSeek might change public perception of AI. Moreover, it has already changed the conversation about what drives progress in artificial intelligence.

DeepSeek has proven to be more efficient and resourceful than other LLM. Beyond its $5.6 million training, DeepSeek has demonstrated something even more important. DeepSeek showed the public how to aim high and work hard to improve their product. It showed how big a role engineering truly plays when developing a large language model. Instead of large investments, it emphasises engineering quality and performance rather than sheer spending.

​The AI industry is evolving continuously, with increasing improvements in architecture, hardware-aware optimisation, reinforcement learning, and systems engineering. Upcoming breakthroughs may be driven not only by larger models and faster hardware, but by smarter algorithms, more efficient architectures, and a fuller understanding of how intelligent systems learn.

DeepSeek shifted the industry's focus from how much compute is used to how effectively it is used. That change in perspective may prove to be its most significant contribution to the future of artificial intelligence.

Despite the emergence of newer, increasingly capable open models such as Qwen3, Kimi K2, and GLM-4.5, DeepSeek's impact on the AI landscape stays significant. These models have developed their own approaches, but many share principles that DeepSeek helped bring into greater focus, including reinforcement learning, Mixture-of-Experts architectures, distillation, inference-time reasoning, and greater computational effectiveness. DeepSeek's importance, therefore, is not simply in its models reaching the frontier, but in how they helped change the industry's thinking about achieving frontier-level performance. Rather than relying primarily on larger models and greater computational resources, researchers are increasingly exploring how improved training, architecture, and optimisation can achieve more with less.

The Real Cost of Training Frontier AI Models: What DeepSeek Revealed and What It Means for the Industry
About 
Propagated AI Team

The Propagated.ai team consists of AI researchers, marketers and product strategists dedicated to helping companies create exceptional digital experiences through the power of artificial intelligence and user-centered design.

View all posts by Propagated Team →
Table of Contents

Free account

One account for the labs, the belt test and the weekly briefings.

Your account

Welcome back.

Not you? Log out