Data that wasn't collected from life but made on purpose to teach the model. It solves two problems at once: there isn't enough real data, and some of it shouldn't be touched. Here's where it helps and where you need to watch out.
Pilots train on a simulator before they get into a real plane. A storm, a failed engine, landing in fog. Nobody waits for the real storm to practice. The simulator is an artificial world built to look enough like the real one. Synthetic data is the simulator for the model.
A model without a single byte of customer data
On September 15, 2026 Salesforce and NVIDIA announced Koa, a reasoning model for working with customer systems (CRM). By NVIDIA's account it was fine-tuned from the open Nemotron 3 Super model on a proprietary synthetic dataset drawn from nearly three decades of CRM deployments across more than 14 industries. In the words of Salesforce's Rohan Kumar, not a single byte of customer data was used. Here synthetic is both a technique and a promise to the customer.
Robotics runs on the same logic. NVIDIA's Cosmos platform generates sensor data for driverless cars in different weather and lighting, so nobody has to wait for real rain on a real road.
Where the catch is
The generator carries its own mistakes. If it describes the world crookedly, the model learns the crooked version and repeats it with confidence. So ask what the synthetic set was checked against. In the Koa announcement the comparison is on Salesforce's own test, CRM Bench: by the company's numbers the model matches or beats leading models and makes three times fewer errors. The test belongs to the company, and that's part of the answer too.