36Kr has learned from multiple independent sources that ByteDance has established a new top-level business unit for artificial intelligence data and security. The unit is on the same organizational level as Seed, Flow, Douyin, and other major business units, and is led by Adam Wang.
It is ByteDance’s latest top-level unit focused on AI, following the creation of Seed and Flow at the end of 2023. Before taking his current role, Wang was head of platform responsibility and head of live streaming at TikTok. “The live streaming business Wang oversaw was once one of TikTok’s biggest sources of revenue, and he had a strong track record within ByteDance,” said a person close to the company.
36Kr contacted ByteDance for comment, but the company had not responded as of publication.
After building out its model and product organizations, ByteDance is now turning its attention more directly to AI data.
According to 36Kr, one precursor to the new unit was a data team established in 2023 by Fu Yue, a member of TikTok’s founding team. The team, which initially numbered around 100 employees, supported overseas businesses such as Dola, the international version of Doubao, and TikTok before taking responsibility for data procurement and quality control for Seed’s model training. It included product managers, data engineers, procurement specialists, quality control and operations staff, and security and compliance personnel.
In addition to that team, the AI data and security unit has consolidated several AI data teams that had previously been scattered across the company, including the group’s DMC platform, and personnel from the AI data platform, or AIDP, under Flow.
Several people familiar with the matter told 36Kr that the integration and formal establishment of the new unit began in early June. Because the consolidation spans multiple departments with overlapping responsibilities, ByteDance is still clarifying the organizational structure and staffing.
Several people close to the new unit said the consolidated organization will be a large data team supporting both foundation models and business teams. Its core function will be to provide cross-modal data services for all of ByteDance’s foundation models. Its responsibilities will span the entire data production process, including setting standards, sourcing and procurement, data synthesis and cleaning, and quality evaluation.
At a Seed all-hands meeting in late July, Zhang Yiming, who has made few public appearances in recent years, said ByteDance would “firmly reject distillation” in its efforts to advance foundation models.
At a ByteDance all-hands meeting on August 5, CEO Liang Rubo also said, “ByteDance’s large language models will firmly remain self-developed. We need to build the fundamentals well, accept falling behind in the short term, and continue optimizing for the long term. Most importantly, we must not lose our direction.”
That emphasis on developing models in-house is likely to make data even more important to ByteDance’s AI efforts.
Beyond ByteDance, the move points to a broader shift across the global AI sector. As the supply of readily available, high-quality public internet data becomes more constrained, model developers are placing greater emphasis on high-quality data generated through real-world activities and workflows.
Inside ByteDance’s data operation
ByteDance has invested heavily in AI data compared with many other major Chinese technology companies.
One reason is Seed’s “no distillation” principle. ByteDance aims for most of its models to reach the global first tier, or even state-of-the-art performance, a goal it does not believe can be achieved through distillation alone. That increases the need to source, synthesize, clean, and evaluate training data internally.
The scale is substantial. According to 36Kr, more than 1,000 people within ByteDance were involved in evaluating model data for Seedance alone. For each Seedance algorithm engineer, more than ten data staff often provided support.
That contrasts with many video AI startups, where internal evaluation teams may number only several dozen people.
Multiple industry practitioners have described the success of Seedance 2.0 as “a victory for data.”
ByteDance has continued to increase its investment in the area.
The increase is visible in its budget. 36Kr previously reported that ByteDance’s data budget for training world models and coding models had already reached an eight-figure USD sum at the beginning of 2026, with instructions that the budget could be increased if necessary.
36Kr also learned that ByteDance’s data teams have adopted a system that assigns parallel teams to the same problem and compares their results. Internally, teams are divided by area, including world models, code, and advanced disciplines. As the economics of models become clearer, ByteDance also requires individual data projects to calculate their return on investment.
Other major technology companies are also increasing their investment in data. Tencent, which stepped up its work on foundation models last year, has recruited from ByteDance’s data teams over the past six months, offering some candidates salaries as high as three times their previous compensation.
Alibaba and Tencent have both increased their data procurement budgets. A person in the industry told 36Kr that major technology companies now employ various data exclusivity strategies, such as setting exclusive periods for datasets or temporarily securing exclusive access to key personnel from suppliers.
Yet the data market in 2026 faces a basic imbalance: AI companies have strong demand, while high-quality data remains in limited supply. That constraint extends from pretraining through post-training.
For ByteDance, a central challenge will be keeping a data organization of around 1,000 people responsive to rapid changes at the frontier of model development. As model capabilities continue to evolve, the types of data that are scarce today may be considerably less valuable six months from now.
Data moves to the center of global model competition
Over the past year, foundation model developers around the world have faced a common constraint: as reasoning approaches converge and gains from architectural changes become harder to achieve, data has taken on greater importance in determining model performance.
Leading overseas foundation model companies are also spending several times more on data than ByteDance and other major Chinese technology companies. External data budgets at leading Silicon Valley model companies are expanding by billions of US dollars a year. Anthropic, for example, allocated more than USD 1 billion to reinforcement learning data in 2025 alone.
That spending has helped fuel rapid growth among data suppliers in Silicon Valley. One example is Mercor, a startup whose annualized revenue rose from USD 500 million last year to USD 2 billion by the middle of this year. About 91% of its revenue came from leading model companies such as OpenAI and Anthropic, while its valuation rose to USD 20 billion. The company was founded only three years ago.
Why has data become so important? One reason is that the pool of readily available, high-quality public internet data has become increasingly constrained.
Over the past three years, the data required for model training has increasingly shifted from the public domain toward private sources. The public internet contains vast amounts of reports and documents, but it captures less of the process through which people produce those results, such as working through ambiguous requests, gathering context, making mistakes, and correcting them.
Beyond coding, for example, general-purpose models still struggle with highly specialized tasks in fields such as medicine, law, and scientific research. As developers seek to improve models on expert-level tasks, they need large amounts of naturally generated “process data.” Such data is often embedded in real-world workflows, including internal code repositories and records generated by employees during their day-to-day work.
As a result, access to proprietary data is becoming an increasingly important competitive front for foundation model companies.
KrASIA features translated and adapted content that was originally published by 36Kr. This article was written by Deng Yongyi for 36Kr.