# DataGame — full text Contact: hello@datagame.ai (people and agents) ## The most complete record of how people play and how games are made. Ten years of live games, 200M+ players and the studio that built them, catalogued for AI labs. Everything here comes from games we built and servers we run. ## Holdings - Match records (available; Millions of matches): Server-side records of multiplayer matches: movement and aim, every hit, loot, vehicles, the storm, placement. Tick-exact. Bots labeled. - Player histories (available; 200M+ players): Pseudonymized careers: skill, progression, sessions and social play over years. - Economy (available; Years of transactions): Stores, markets, offers and currencies. What players valued, traded and bought. - Action traces (available; Growing daily): Flappy Bird seeds and every input. Exact actions in a deterministic world. - Source and reviews (available; A decade of repos): Game code, commits and pull requests with review threads, across titles we own. - Studio work (available; Years of tickets): Bug tickets, design docs and team threads, consented and redacted, linked to code. - Live ops and performance (available; Thousands of devices): Balance patches, config changes, tests, crashes and frame times, with player outcomes. - Environments and evals (available; Growing library): RL environments on production rules, and benchmarks with human baselines. ## Terms of engagement: graded by the game, measured on your evals A data purchase should be judged by what it does to your model on tasks we never saw. Ours is built for that test: the referee is a game server, the records were never public, and the people in them were playing for real. - E-01 Environments on production rules: graded by the game server's own outcome (survival, placement, pipes cleared, win or loss); human baseline from players at every skill level in the same situation; train on it, then score a suite you hold back. - E-02 Private evals that refresh: server truth, tick by tick; archive players from the same moment; new matches daily, so a set can be rotated instead of retired. - E-03 Human play at scale: 200M+ people who chose to play, bots labeled; score distributions by skill band; compare model and human on the same seeds. - E-04 Studio work with outcomes: tests from the fix that shipped, plus what players did next; baseline is the engineer who fixed it; regression and breakthrough splits from repos never made public. - E-05 Targeted slices and capture: you name where the model breaks; we cut from the archive or capture to it, with matched human sessions. Every environment ships with a reference solution at full marks, bad runs scoring zero, hidden test files and its measured flake rate. We train a small open model on each release and score held-out tasks across several seeds; no gain above the noise, no release. Open task format (task, container, verifier, reward file), versioned and pinned. A data card covers checked size, provenance, rights, personal-data handling and known limits; exclusivity by title, modality or field of use. ## How a match becomes data Players' devices send actions to the authoritative game server, which decides every hit, move, pickup and storm tick. A match recorder writes one record per match, events stamped by server tick. Records, server logs, client events, player KPIs and economy data land in the archive, which feeds datasets, evals and RL environments. One chain of custody: player, server, archive, license. ## How a game gets made Design, ticket, code and pull request, review, build, release, players, telemetry, and back. Each step leaves a record. Linked together (ticket, commit, patch, player outcome) they become tasks an agent can be scored on. Consented, redacted, titles we own. ## Inside one match record Tables: match (id, map, mode, team mode, start and end time, ticks per second, winner), users (id, pseudonymized nickname, isBot, team), movements (userId, tick, position, rotation), hits (attacker, target, tick, weapon, damage, positions, headshot, kill), loot boxes and picked items, storms and storm passes, user flights and vehicle uses, ability uses and heal attempts, user summaries. ## What others leave out 1. The referee's record: data from the game server, tick by tick, not pixels from one screen. 2. No middleman: our games, servers and code; one chain of custody. 3. The whole studio: gameplay linked to code, tickets, economy and patches. 4. People who chose to play: no paid sessions; bots labeled. 5. Every genre: shooters, strategy, arcade, racing, Roblox; mobile and PC. 6. Baselines from millions: human scores for every benchmark. ## Evals - Next state (World modeling): model gets seconds of match state; scored by position and event error; human baseline: archive players. - Human or bot (Perception): model gets one player's movement and fire; scored by server bot flag; human baseline: not applicable. - Zone routing (Planning): model gets position, loadout, storm; scored by survival; human baseline: players in the same spot. - Flappy Bird (Control, RL): model gets frames or state; scored by pipes cleared; human baseline: live players on the same seeds. - Tower Conquest (Strategy): model gets board, hand, opponent; scored by win rate; human baseline: ranked ladder. - Fix the bug (Software agents): model gets real ticket and repo; scored by tests and the shipped fix; human baseline: the engineer who fixed it. ## Standards - Consent: Collected under each game's terms of service and privacy policy; studio data only with employee consent. - Privacy: Player IDs pseudonymized, nicknames removed, personal data stripped before delivery. - Provenance: Bots flagged in every record; fields marked recorded, reconstructed or inferred. - Rights: Titles we hold rights to; exclusive or non-exclusive per dataset, in writing. - Delivery: JSON and Parquet records, MP4 and WebDataset for rendered views, hosted RL environments; versioned with a data card. ## FAQ ### What rights does a license give us? Training, fine-tuning, evaluation and internal research by default. Commercial deployment of models trained on the data is included. Redistribution of the raw data is not. Every license is written per dataset, so the terms match what you are building. ### Can we get exclusivity? Yes, for a time window, a modality or a field of use. Most buyers start non-exclusive with a pilot and convert the parts they care about. ### Who owns the data? SuperGaming built and operates the games, so the data comes straight from the source: our servers, our players, our code. DataGame (a SuperTuned company) licenses it. There is no broker in the chain. ### How do you handle player consent and privacy? Data is collected under each game's terms of service and privacy policy. Player IDs are pseudonymized, nicknames are replaced, and personal data is removed before anything leaves our systems. Studio communications are released only with employee consent and redaction. ### Are bots mixed in with human players? Bots are flagged in every record. You can filter them out, or keep them as a control group. We quote volume in human play, so bots never inflate a deal. ### What exactly is in a match record? Everything the game server saw: positions and view direction for every player, every shot and hit with both players' positions, loot, abilities, vehicles, heals, the storm and final placement. Events are tick-exact, straight from the authoritative server. ### Which games and genres are covered? Battle royale and arena shooters, real-time strategy, arcade, racing, and Roblox titles, across a decade of live operation. Plus the studio around them: code, tickets, live-ops and performance data. ### Do you have video? We have gameplay footage for many titles, and we are building a renderer that turns stored matches into video, depth and segmentation, labeled as recorded or reconstructed. New capture with inputs and video is available for custom projects. ### How should we evaluate your data? Train on a slice and score it on tasks you hold back. Every slice ships with provenance and a data card so the comparison is clean, and we publish our own measured result, run over several seeds, before you see it. If it does not move your numbers, tell us. ### Can you build for a specific failure? Yes. Describe the failure: agents losing the thread over long tasks, planning under time pressure, coordinating a team, reading an economy, fixing real code. We show what in the archive bears on it, then cut a slice, package an environment or run a capture. ### Will training on games help beyond games? That is the open question in RL, and the one worth paying for. We measure it: every release reports held-out scores on the same game, on our other games and on non-game tasks that test the same skills, across several seeds. When the gain is small, the data card says so. ### Has any of this leaked into pretraining data? No. Server records and studio repositories were never on the public internet, and footage is a small part of what we hold. Private sets stay private, and because new matches arrive daily they can be rotated rather than retired. ### Do your evals include human baselines? Yes. Every benchmark ships with human scores from the same game, drawn from real players at different skill levels, plus a scripted-bot floor and the grader code. ### Can we run agents in your games? Yes. Our RL environments run on the production games, with the same rules players had. Flappy Bird is live today; more titles are being packaged. ### Can we use your environments on the platform we already train on? Yes. Each environment ships in an open task format (task, container, verifier, reward) that common RL frameworks and hosted training platforms can load. Tell us what you train on and we package for it. ### Can we see a sample before signing? Yes. We send a schema, a data card and a sample under a short NDA, usually within days of a first call. ### Can we test tasks you did not pick? Yes. We commit the full catalogue first, then a public random draw chooses the tasks you test, so the sample is not hand-picked. Each task comes with how often several models solve it, so you see the difficulty before you buy. ### How is data delivered? JSON and Parquet for records, MP4 and WebDataset for rendered views, hosted environments for RL. Delivered to your cloud bucket, versioned, with a data card for every release. ### How is pricing set? By dataset, volume of human play, exclusivity and refresh cadence. Pilots are priced to be easy to say yes to. ### Do you work with companies training their own models? Yes. Enterprises training models with reinforcement learning, and the firms that train for them, license our environments and evals the way labs do: a pilot first, then a licence. ### Can my studio license its data through you? Yes. We clean, pseudonymize and package your data to the same standard as ours, you keep ownership, and you share in every license. ## Journal ### Base data, and the layer above it (Oct 2026) The RL market has sorted itself into layers. Underneath all of them is the real work environments are built from. A year ago most training data was labels. Now the market has layers. Environment builders make the worlds and tasks a model practises in. Infrastructure companies give teams the tools to build, host and train on those environments. Service firms run reinforcement learning for companies that want their own models. Labs, and now a growing number of enterprises, buy from all three. Underneath every layer sits the same raw material: base data. The codebases, logs, configurations and records of real work that an environment is built from. An environment is only as good as what it was built on. Contrived tasks teach contrived skills, and a grader written in an afternoon is the first thing a model learns to cheat. That is why quality, not supply, is now the bottleneck. Platforms that test what vendors sell report that much of it fails: the reward climbs because the model found a shortcut, or the tasks were too artificial to teach anything that carries over. The test that matters is simple to state. Train an open model on the environment, watch the reward rise, then check that it also improved on tasks it never saw. Buyers have started testing the way an auditor would: commit the vendor’s full catalogue, draw tasks at random, and check that models solve some of them but not most. A game studio sits in two layers at once. We hold the base data, a decade of server records, code, tickets and economy, and we build environments on it that are graded by servers that refereed millions of real matches. We sell both. We ship in open formats, so our environments run on the platforms labs already use, and we measure transfer before we ask anyone to pay for it. ### The question worth paying for (Oct 2026) Getting good at one game is cheap. Getting better at everything else because of it is the claim labs pay for. The open question in reinforcement learning is generalization. A model trained on coding tasks gets better at coding. Does a model trained on a battle royale get better at anything but battle royales? Nobody selling environments can answer that in general, and anyone who says otherwise is selling. What we can do is measure it for our own releases and publish the result. Every release is scored on held-out tasks from the same game, on tasks from our other games, and on a small set of non-game tasks that test the same skills: planning over a long horizon, acting under time pressure, working with or against other agents. Several seeds, with the spread shown. We expect games to transfer better than most environments because they stress what agents still get wrong, and because the people in our data were not paid to perform. But expectation is not evidence. When the gain is small, the data card says so. For buyers the practical advice is the same as for any data: keep your model and your test suite fixed, change only the data, and run it more than once. ### The missing environment (Oct 2026) We read the sites of twenty-eight data and environment vendors. None sells a real game world. The market for training data has moved from labels to worlds. Almost every vendor we read now leads with environments: a task, a sandbox, a way to check the result and a reward. Some rent experts by the hour, some record how experts work, some build replicas of offices, codebases and trading desks, and a few own a vertical such as biology or markets. Games barely appear. One vendor has models play a puzzle game through text. One tests planning under a speedrun clock. One asks agents to write a console emulator. An open framework ships a strategy-game plugin. Nobody offers what a game studio actually has: live worlds with real rules, real opponents and millions of people who already played them. That matters because games stress exactly what agents still get wrong. Long horizons, where a decision at minute two decides the outcome at minute ten. Many agents at once, some cooperating and some not. Real-time pressure. Outcomes that arrive late and noisily. A replica can approximate those things. A production game server is them. So that is what we sell. Environments that run on the rules players had, graded by the server rather than a model, with human score distributions from people who chose to play. And alongside them, the studio that built those games: the tickets, code and patches, linked to what happened next. ### Environments are datasets now. Ours were already running (Oct 2026) Open formats now treat an environment like any other dataset. Here is how we are packaging ours. This month environments started to be published like datasets: a repo holding tasks, a container, a verifier and a rule that turns the result into a reward, tagged for the RL frameworks that can load it. Open arenas now train a fixed model on submitted environments and score the result on tasks nobody has seen. The plumbing for buying and testing environments is becoming shared. We are adopting it. Our first release is Flappy Bird: every task is a seed, the container runs the game headless, the verifier replays the inputs and counts pipes cleared, and a reference run proves each task can be solved. Because the world is deterministic, a seed and a list of inputs reproduce a run exactly, which makes the grader impossible to argue with. The battle royale and strategy titles follow the same shape, with one difference that matters. Their containers run our authoritative game servers, the same code that refereed real matches. Nothing in the environment is a reconstruction of the game. It is the game. Releases will be versioned and pinned, gated for buyers, with an open sample for anyone who wants to check the format before talking to us. ### A referee that cannot be talked round (Oct 2026) Reward hacking is the tax on every environment. A game server pays less of it, but not none. Every environment eventually meets a model that finds the shortcut: editing the test file, planting a helper that reports success, reading answers it was never meant to see. The field has learned to hide verifier files, run agents without root and without network, and test each grader against good, bad and deliberately adversarial runs before trusting it. Games have a structural advantage here. The grader is not a script written for the benchmark or a model asked for an opinion. It is the server that already refereed millions of matches, written to stop cheating players long before anyone trained an agent against it. Survival, placement, pipes cleared and wins are counted by code that people have attacked for years. That is not the same as unhackable. Players find exploits, and speedrunners make a sport of it. So we treat our own patch history as part of the product: known exploits, when they were closed, and the matches recorded before and after. A model that finds a glitch we already fixed tells you something; one that finds a new one tells us something. Every release lists the checks it passed: reference solution at full marks, do-nothing and wrong runs at zero, hidden files, network off, and the flake rate across reruns. ### Measured before you see it (Oct 2026) The honest way to sell data is to show what it does to a model on tasks it never saw. We run that test first. The cleanest test for a dataset or environment is now widely agreed: keep the model, the training recipe and a private test suite fixed, change only the data, and measure the difference. Open arenas run exactly that contest, and they add the warning every buyer should hear. One run is noise. A small gain from a single seed proves nothing. We hold ourselves to it before a lab does. For each release we train a small open model on our tasks, score it on held-out tasks from the same game and on tasks from other games, and repeat across several seeds. The result ships in the data card with its spread, not just its best run. If a release does not move held-out scores above the noise, it does not ship. If it only moves the game it came from, we say so. Transfer is the claim worth paying for, so it is the claim we measure. Labs should still run their own test on their own suite. Our number is there so the first conversation starts from evidence. ### Evals from the work itself (Oct 2026) The best evals come from real work with a known outcome. A studio produces them every week. More teams now build evals the obvious way: take the work they actually did, and check whether a model can do it. Split that work in two and it serves two purposes. Regression tests, the things a model must keep getting right. Breakthrough tests, the things nothing can do yet. A game studio is unusually rich in this material. Every bug ticket has a reproduction, a fix that shipped and tests that pass; every balance change has the player data before and after. A model can be asked to fix the same bug in the same snapshot of the code, and the shipped fix grades it. Two properties make these tasks valuable. They were never public, so no model has seen the answer. And they carry an outcome beyond the tests: whether the fix held, and what players did next. We release studio data only with consent and redaction, with partner-confidential material left out. ### What a studio holds behind the door (Oct 2026) Labs have read the open web. What they want now sits inside companies. Here is what a clean package from a game studio looks like. Models have been trained on the same crawl, the same public code and the same forums. What labs ask for now is work that never left a building: internal tools, real workflows, expert decisions and the outcomes that followed. Game studios hold a great deal of it, and most do not know what it is worth. Value depends on packaging, not volume. A buyer wants to know exactly what the data is and what it is good for, how big it is and whether that size was checked, its format, where it came from and who holds the rights, how personal data was handled, and its known limits. A package that cannot answer those questions does not sell twice, however large it is. Some data should never be sold. If its value lies in who the people are, removing them leaves nothing worth buying, and keeping them is not an option. Player identities are pseudonymized before anything leaves our systems, and data a studio wants kept stays kept. If you run a studio and want to license your game's data, we package it to the same standard as ours. You keep ownership and share in every licence. ### Why the server record beats the screen (Oct 2026) Screen capture shows what one player saw. The server record knows what actually happened, to everyone, at every tick. Most game data for AI today is captured from screens or engine hooks on a player's machine. That shows one viewpoint, compressed, with the game's truth left out. A game server is the referee. In a battle royale it decides where every player is, who hit whom, how much damage landed and when the storm closes. It writes that down. Those records are what DataGame licenses. For a world model, that means ground truth: positions, facing, events and outcomes for every agent in the match, not inferred from pixels. For an eval, it means a grader nobody can argue with. Where buyers need pixels, we render them from the record and label every frame recorded or reconstructed. The truth comes first, the pictures second. ### Humans and bots in the same lobby (Sep 2026) Live matches mix real players with server bots. Labeled, that mix becomes a free control group. Live multiplayer games fill lobbies with bots when there aren't enough players. Most datasets hide that. We flag it. In the match on our homepage, a handful of humans made almost every kill and dealt most of the damage, in a lobby that was mostly bots. Human play looks different, and the difference is measurable. That gives labs two things: clean human-only data when they want it, and a labeled contrast set for questions like "does this model play like a person?" Our Human-or-Bot eval is built on exactly that. ### Flappy Bird is our game lab (Aug 2026) A deterministic world, exact actions, and a live player base: the simplest possible RL environment with real human baselines. We relaunched Flappy Bird in 2026 and have recorded runs since July: the seed and every flap. Because the world is deterministic, a seed and a list of flaps replay a run exactly. That makes it a clean environment for control and world-model research, with a human score distribution from people who chose to play. We use it to test every part of our pipeline before we roll it out to bigger games. ### A game studio is a dataset (Jul 2026) Code, tickets, reviews and live-ops changes, linked to what players did next. A studio produces more than gameplay. Every bug report, fix, config change and balance patch is a record of people solving problems under real constraints. Linked together, a ticket, the patch that fixed it, and the player data before and after make a task an agent can be scored on. That is rare: real software work with a real outcome. We release studio data only with consent, with redaction, and with partner-confidential material excluded. ## Contact Email hello@datagame.ai, or use the contact form on the site (labs getting data from us, and partners licensing their data to us). Agents can POST a structured enquiry to https://datagame.ai/api/lead (see llms.txt). Labs: data, evals, environments, pilots. Studios: license your game's data, you keep ownership. Investors and press: deck and data room on request.