Following Seed and Flow, ByteDance Establishes Another AI Tier-1 Department, Targeting "Data"
I'm LongbridgeAI, I can summarize articles.ByteDance has recently established a new Tier-1 AI department, "AI Data & Security," led by Wang Yinglei. This department integrates teams such as Global Data, the Group's Data Middle Platform (DMC), and AIDP under Flow, aiming to provide cross-modal data services for all of ByteDance's large models, covering the entire data production lifecycle. This move marks another core department established around AI business following Seed and Flow, with ongoing structural reorganization and personnel optimization currently underway
"Intelligence Emergence" has learned from multiple independent sources that ByteDance has recently established a new Tier-1 department—AI Data & Security—parallel to departments such as Seed, Flow, and Douyin. The head of the department is Wang Yinglei (Adam Wang).
This is another Tier-1 department established by ByteDance around its AI business, following the creation of two AI Tier-1 departments, Seed and Flow, in late 2023. Wang Yinglei previously served as the Head of Platform Responsibility and Head of Live Streaming at TikTok. "The live streaming business managed by Wang Yinglei was once one of TikTok's primary revenue sources, and he has achieved significant success within ByteDance," said a source close to the company.
"Intelligence Emergence" sought confirmation from ByteDance regarding the above information, but there was no response at the time of publication.
After focusing on models and products, ByteDance is now directly targeting AI data.
According to "Intelligence Emergence," one of the predecessors of this new department was the Global Data team, established in 2023 by Fu Yue (nickname: Yuyi), a founding member of TikTok. The original team size was around 100 people. This lean team initially served international businesses such as Dola (the overseas version of Doubao) and TikTok, and later became responsible for data procurement and quality control for Seed's model training. The team includes various functions such as product management, data engineering, procurement, quality inspection operations, and security compliance.
In addition to Global Data, the AI Data & Security department has integrated several previously dispersed AI data departments, including the Group's Data Middle Platform (DMC) and personnel from teams such as the AI Data Platform (AIDP) under Flow.
Several informed sources told "Intelligence Emergence" that the integration and official establishment of this department began in early June 2026. Due to the involvement of multiple departments with overlapping businesses, ByteDance is currently continuing to streamline its structure and optimize personnel.
Several individuals close to the department told us that the integrated new department will be a massive data team spanning from foundation models to business units. Its core function is to provide cross-modal data services for all of ByteDance's large models. The team's responsibilities also cover the entire data production lifecycle: standard setting, sourcing and procurement, synthetic cleaning, and quality evaluation.
At Seed's all-hands meeting in late July, Zhang Yiming, who had not appeared publicly for some time, stated: ByteDance is "firmly against distillation" in its large model efforts. At ByteDance's all-hands meeting on August 5, CEO Liang Rubo also stated: "ByteDance's large language models will remain committed to in-house R&D, strengthen fundamentals, accept short-term lag, persist in long-term optimization, and most importantly, ensure that actions do not deviate." These decisions will inevitably lead to an increasing importance of data in ByteDance's large model endeavors.
Looking beyond ByteDance, we also see structural changes occurring in the global AI industry: as public internet data becomes exhausted, the core variable in model competition is shifting from algorithms and computing power to high-quality data from the real world.
ByteDance's Massive Data Empire
ByteDance has been the most resolute among domestic tech giants in its investment in AI data.
An important reason is that Seed established the principle of "no distillation" from its inception. ByteDance's goal for almost all its models is to reach the global first tier or even State-of-the-Art (SOTA), which is difficult to achieve through distillation. If all data must be synthesized, purchased, and cleaned internally, it inevitably requires a massive team to support these efforts.
How large is this team? "Intelligence Emergence" previously reported exclusively that within ByteDance, the team dedicated solely to model data evaluation for Seedance numbers over a thousand people. Behind every Seedance algorithm engineer, there are often more than ten data colleagues providing support.
In contrast, for leading startups in the video sector, an internal evaluation team of dozens is already considered a significant investment. The success of Seedance 2.0 has been described by many practitioners as "a victory of data."
This has further strengthened ByteDance's commitment to data investment.
First, at the budget level. "Intelligence Emergence" previously reported exclusively that for the training of world models and coding models, ByteDance's data budget in early 2026 had already exceeded the tens of millions of dollars level, and "if deemed insufficient, additional budget can be added at any time."
"Intelligence Emergence" also learned that ByteDance's data team has adopted a horse-racing mechanism, with teams divided by direction—world models, code, difficult disciplines, etc. Against the backdrop of increasingly clear commercial logic for models, each of ByteDance's data projects also needs to calculate ROI.
Beyond ByteDance, other major tech companies are strengthening their data investments. Tencent, which truly began to exert effort in large models last year, has frequently poached talent from ByteDance's data team in the past six months, offering salaries up to three times higher.
Alibaba and Tencent have both seen significant increases in their data procurement budgets and are continuously increasing their investments. A data industry insider told "Intelligence Emergence" that major tech companies currently have certain data exclusivity strategies, such as setting exclusive periods for datasets or temporarily buying out key personnel from suppliers.
"Data is the most decisive variable in the current competition for large models." This has become a consensus across almost all star sub-sectors, including large language models, embodied intelligence, world models, and AI4S.
However, the current state of the data market in 2026 is that AI companies have strong demand, facing a bottleneck of scarce high-quality data supply from pre-training to post-training.
For ByteDance, the next challenge is how to enable a data organization of thousands of people to respond quickly to changes at the forefront of model development. After all, during a phase where the boundaries of model capabilities are changing rapidly, the scarcest data today may lose its value in half a year.
Global Model Competition Turns into a Data Arms Race
Over the past year, global large model efforts have faced a trend: under the premise of slow convergence in reasoning paradigms, the capability improvements brought by algorithmic architectures are slowing down. Data has become the most important variable determining the upper limit of model capabilities, if not the only one.
Compared to Chinese tech giants like ByteDance, overseas large model companies are investing several times more in data. The external data budgets of top Silicon Valley large model companies are expanding rapidly by billions of dollars annually. For example, Anthropic allocated a budget of over $1 billion for RL data in 2025 alone.
This has made data one of the fastest-growing sectors in Silicon Valley in 2026. A typical example is the Silicon Valley star company Mercor, whose annualized revenue grew from $500 million last year to $2 billion by mid-year this year, with 91% coming from top model companies like OpenAI and Anthropic. Its valuation soared to $20 billion, despite the company being only three years old.
Why has data suddenly become so important? The fundamental reason is that public data on the internet has been almost exhausted.
Over the past three years, the data required for model training has shifted from the public domain to the private domain. While there are abundant reports and documents on the public internet, they lack the process of how humans produce these results, such as how to understand ambiguous intents, search for context, make mistakes, and correct them.
For instance, beyond coding, there are high-end tasks in fields such as medicine, law, and scientific research, which general models have not yet solved well. When models need to improve their capabilities to work like real experts, they require large amounts of naturally occurring "process data." This data is often hidden within real workflows, such as code repositories accumulated within enterprises and traces left by employees' daily operations.
The data required for AI is shifting from the public domain to the private domain. And the collection of large amounts of private domain data has become a new point of contention for large model companies.
Risk Disclosure and Disclaimer
The market involves risks; investment should be approached with caution. This article does not constitute personal investment advice, nor does it take into account the specific investment objectives, financial status, or needs of individual users. Users should consider whether any opinions, views, or conclusions in this article align with their specific circumstances. Investors bear full responsibility for their own decisions.
