DeepSeek is an AI team based in Hangzhou, China. In January 2025, it shook the market with R1, an open model billed as having been trained at unusually low cost. Nvidia lost $589 billion in market value in a single day, an event that came to be described by some as the “DeepSeek moment.” More than a year later, the company introduced its 2026 flagship V4 line and completed roughly $7 billion in funding. The story around DeepSeek now spans its origins, model roadmap, product access, licensing terms, and a growing list of privacy and geopolitical disputes.
From High-Flyer to DeepSeek
DeepSeek was established in July 2023 as an AI research team spun out of High-Flyer, a Chinese quantitative hedge fund. High-Flyer itself was founded in 2015 and built its own GPU clusters for quantitative trading, which gave it access to the compute infrastructure needed for large-model training.
Founder and CEO Liang Wenfeng also serves as CEO of High-Flyer. Through shareholding entities, he controls about 84% of DeepSeek. The company is headquartered in Hangzhou and has kept the team intentionally lean and research-focused. In 2025, it had only about 160 people. Unlike many AI companies, DeepSeek did not put near-term commercialization at the center of its strategy and instead aimed at artificial general intelligence, or AGI, from the outset.
After the first financing round in 2026, Liang’s net worth was estimated to have doubled to about $36 billion, and he was at one point described as the world’s richest AI founder.
The January 2025 “DeepSeek moment”
On Jan. 20, 2025, DeepSeek released R1, a model centered on reasoning performance. It was positioned against OpenAI’s o1 while claiming far lower training compute than rivals. The market reaction followed quickly. On Jan. 27, Nvidia lost $589 billion in market value in a single session, the largest one-day drop ever recorded by a single U.S. listed company, while its share price fell about 17%. On the same day, DeepSeek overtook ChatGPT to become the No. 1 free iOS app in the United States.
One of the most cited figures tied to DeepSeek is the claim that training cost only $5.6 million. In the source material, that number comes from the technical report for DeepSeek V3 and refers only to the compute cost of the final formal training run, using about 2,048 H800 GPUs over roughly 55 days. It does not include earlier research, experiments, data costs or labor. Semiconductor research firm SemiAnalysis estimated that DeepSeek’s cumulative GPU capital spending may have reached about $500 million. The smaller figure, then, is presented as a narrow marginal compute cost rather than the company’s full investment.
How the model line evolved from R1 to V4
DeepSeek’s model family broadly follows two tracks: the V line for general-purpose models and the R line for reasoning-focused models. From V2 and V3 in 2024 to R1 in January 2025, the company moved step by step toward folding “thinking” capability into its main line.
V3.1, released in August 2025, was the first to adopt a hybrid thinking architecture. V3.2, released later that year, added sparse attention, or DSA, to improve long-context efficiency. Industry chatter at one point suggested that a standalone R2 would follow, but that model was not released as a separate product. Instead, reasoning capability was merged into the hybrid route running from V3.1 through V4, according to the source’s reference to DeepSeek’s official update log.
In 2026, those strands converged in V4. The new flagship includes an optional thinking mode, effectively absorbing the reasoning work previously associated with R1. DeepSeek unveiled a preview of V4 on April 24, splitting the family into the more powerful V4-Pro and the lighter V4-Flash. Both support a context window of 1 million tokens, enough to ingest a full codebase or a lengthy document set in one pass. On July 31, DeepSeek released the production version of V4-Flash, strengthened its agent capabilities and cut API pricing again. As of early August, the larger V4-Pro remained in preview.
Efficiency became one of the main talking points around V4. According to MIT Technology Review, V4-Pro requires only about 27% of the compute and 10% of the memory used by the previous V3.2 generation. V4-Flash pushes that down to 10% of compute and 7% of memory. On performance, DeepSeek said V4-Pro stands alongside top closed models such as Anthropic’s Claude Opus 4.6, OpenAI’s GPT-5.4 and Google’s Gemini 3.1. The source explicitly notes that these benchmarks were published by DeepSeek itself and should be treated as vendor claims. V4 was also optimized for Chinese domestic chips including Huawei Ascend.
How people use DeepSeek
For general users, the simplest route is the official web product at chat.deepseek.com or the DeepSeek mobile app. The interface is similar to ChatGPT, supports free chat, and offers a deep thinking mode.
Developers usually access the models through the API. Based on the company’s published pricing cited in the source, V4-Flash costs about $0.14 per million input tokens and $0.28 per million output tokens. V4-Pro costs about $0.44 for input and $0.87 for output, with lower prices when cache hits apply. The article says those rates often come in at only a fraction of comparable closed U.S. models, which helps explain the model’s fast uptake among developers.
Because DeepSeek makes model weights available, advanced users can also run the models locally, including through Ollama on personal computers or servers, which keeps data off the cloud. The full V4 models are large and require high-end hardware, but DeepSeek has also released smaller distilled versions that make local experimentation possible on more modest machines.
Open weights, MIT licensing and distilled models
Since January 2025, DeepSeek’s model weights have been released under the permissive MIT license, allowing commercial use and retraining. That licensing stance is one reason the models have spread so widely.
The source also draws a distinction between open-weight and fully open-source. DeepSeek has published model weights and technical reports, but not the training data or the complete training pipeline. To lower the hardware barrier, the company distilled R1’s reasoning ability into smaller models ranging from 1.5B to 70B parameters. Those models are based on Qwen and Llama and are suited to local machines and edge devices.
2026 financing, a paused second round and Liang’s compute remarks
In June 2026, DeepSeek completed its first external fundraising round at roughly $7 billion. Investors included Tencent, battery maker CATL and the state-backed National AI Industry Investment Fund. The company’s valuation reached about 350 billion yuan.
It then moved into a second round, targeting at least 10 billion yuan in new capital, with the pre-money valuation at one point reportedly reaching about 480 billion yuan. That round was paused in July. The immediate trigger was the leak of remarks Liang Wenfeng made at an investor meeting. He attributed the gap between Chinese AI and the U.S. to compute resources rather than talent. According to a meeting record obtained by Yicai, Liang said, “Our biggest gap with the United States lies in resources.”
Liang also said DeepSeek had compute equivalent to about 20,000 Nvidia H-series GPUs and was working closely with Huawei, while using in-house tools such as TileLang to get around Nvidia’s CUDA ecosystem. Management was described as frustrated by the leak of non-public comments, and investors were verbally told to pause signing. The article says the fundraising process may still restart later. It also says the company is preparing for a possible IPO filing as early as the end of 2026.
The main controversies around DeepSeek
The source groups the disputes around DeepSeek into three main categories.
- Data privacy. DeepSeek’s privacy policy says user data is stored on servers in China. That includes prompts, IP addresses, uploaded files and even typing patterns. The article adds that Chinese authorities have broad legal powers to request data from domestic companies.
- Content controls. Hosted versions of the model may refuse to answer, or may respond in ways aligned with official Chinese positions, on politically sensitive topics such as Tiananmen and Taiwan.
- Government restrictions. Italy removed the service in January 2025 over data concerns. In the United States, the National Defense Authorization Act and related bills were cited as barring use on federal devices, covering agencies including NASA, the Navy, the Pentagon and the Commerce Department. Taiwan’s digital affairs authority also required government departments, state-owned enterprises and public schools not to use DeepSeek in late January 2025. Australia, South Korea and India were also described as having similar restrictions.
The article also notes two additional allegations. In January 2025, OpenAI and Microsoft said they were investigating whether groups connected to DeepSeek had improperly obtained data by “distilling” through the OpenAI API. In 2026, reports also appeared saying Anthropic accused DeepSeek of extensively scraping Claude conversations. The source says these remain allegations without a public conclusion, but they add to ongoing questions over the legitimacy of the company’s data sources.
Positioning against Claude, ChatGPT and Gemini
Against closed flagship models from OpenAI, Anthropic and Google, DeepSeek’s positioning is presented in straightforward terms: open weights, low prices, and benchmark claims in coding, mathematics and agent tasks that aim to show near-frontier performance at a much lower cost.
The article again stresses that most of the performance comparisons with top-tier models are based on DeepSeek’s own published results and should not be treated as independent validation. For developers, DeepSeek may fit cost-sensitive work, self-hosted deployments, or cases where avoiding dependency on a single vendor matters. For data sets with privacy or compliance sensitivity, the article says the fact that hosted data is stored in China, together with government-use restrictions in several jurisdictions, needs to be part of the decision.
The source’s bottom line is that DeepSeek has become one of the most consequential options in the open-weight camp, but where it should be used and what kinds of data it should handle are separate questions.
Questions highlighted in the source
Which country is DeepSeek from?
It is a Chinese company headquartered in Hangzhou and was spun out of High-Flyer in July 2023.
Can DeepSeek be used for free?
Yes. The web version and mobile app offer free chat. API access is usage-based and priced on the lower end, and model weights can also be downloaded for local use.
What is the latest DeepSeek model?
As presented in the article, the latest family is V4 in 2026. The lighter V4-Flash entered production at the end of July, while the larger V4-Pro was still in preview in early August.
Did DeepSeek really train the model for only $5.6 million?
That number refers to the compute cost of the final formal training run for V3 and does not include R&D, experiments or labor. SemiAnalysis estimated GPU capital spending could be around $500 million.
Is DeepSeek safe to use, and is the data stored in China?
The hosted version stores data on servers in China, and several governments have cited that issue in restricting official use. For users with stronger privacy concerns, the source points to local deployment with open weights as the alternative.

