Back
Someone Scraped 5.6 Billion TikTok Videos and Put the Data on Hugging Face for Free
News

Someone Scraped 5.6 Billion TikTok Videos and Put the Data on Hugging Face for Free

A developer posted metadata for about 5.6 billion public TikTok videos, scraped through the app's private API. The free dataset doubles as a storefront.

10/7/20265 min read76 views

What happened

A dataset containing metadata for roughly 5.6 billion public TikTok videos has appeared on Hugging Face. It was posted by a developer who collected the data through the app's private API — a method explicitly prohibited by TikTok's terms of service. The dataset is distributed for free, yet according to Decrypt it also acts as a storefront: the author appears to monetize extended or updated data.

Why it matters for the market

The 5.6 billion figure is not just a technical milestone. For teams in digital marketing, traffic arbitrage, and social analytics, metadata at this scale opens possibilities that previously required in-house scraping or expensive third-party tools. Such metadata typically covers video-level attributes (identifiers, publishing parameters, engagement metrics) rather than the content itself, and it powers trend discovery, creative evaluation, and audience pattern analysis.

At the same time, this is a story about legal and reputational risk. Scraping via a private API breaks platform rules, so any business planning to use such a dataset commercially must assess the legal exposure, including data protection concerns if indirect identifiers are present.

Market context

  • Major platforms have spent years tightening API access and fighting scraping, making large public dataset leaks more visible events.
  • For arbitrage and media buying, the value lies in quickly spotting viral formats and understanding what drives engagement.
  • Demand for grey and black data sources keeps growing, raising risks for those who build durable processes on them.

Editorial take

The appearance of this dataset signals that social data is becoming a commodity in its own right, even when nominally free. For digital marketing and arbitrage professionals, it is a reminder to revisit sourcing policies. Relying solely on grey datasets is a short game: access can be cut and regulators can ask uncomfortable questions. The more resilient approach combines public, legitimate APIs with proprietary analytics pipelines. Treating such datasets as an exploratory tool is sensible; building long-term infrastructure on them is not.

Share this article

Get the best affiliate marketing jobs first

Subscribe to our Telegram channel

Stay connected in Telegram

Follow the latest posts in Telegram

@tiktok_arbitrajFresh updates
Join @tiktok_arbitraj

Looking for talent? Post a job

18,000+ Telegram subscribers, 24,000+ jobs on the platform. Posting from $39.