AI Training Data

Automotive training data for AI and machine learning

226M+ listing records and 43M+ real auction sales, VIN-keyed and cross-platform deduplicated — the price signal you need to train valuation, pricing, and demand models across new and used vehicles.

226M+
listing records
43M+
auction sales
2015
earliest history
10+
platforms
Parquet · CSV · JSON
formats

What's In The Data

Every field your model needs, keyed on VIN

Normalized, VIN-keyed records you can drop straight into a feature store — listing prices, daily price history, mileage, condition, and source provenance on every row.

vin Primary join key — 17-character VIN linking every dataset.
price Asking price in USD at time of listing.
price_history Daily price snapshots captured while the listing was live.
source Platform of origin — a 10+ value enum for source weighting.
condition One of new, used, or certified_pre_owned. Always populated.
mileage Odometer reading at the time the listing was observed.
{
  "vin":           "1FTFW1E84MFA12345",
  "first_seen_at": "2024-08-12T14:22:08Z",
  "last_seen_at":  "2026-04-29T09:14:51Z",
  "year":          2022,
  "make":          "Ford",
  "model":         "F-150",
  "trim":          "Lariat SuperCrew 4x4",
  "mileage":       38210,
  "price":         42950,
  "condition":     "used",
  "location":      { "zip": "30309", "state": "GA" },
  "source":        "cargurus",
  "price_history": [ 7 snapshots ... ]
}

Who Uses It

Who trains on this data, and why

Real teams, real models — built on listings and auction prices in one VIN keyspace.

For valuation startups

Train an AVM from scratch

ML engineers train automated valuation models on 226M+ listings plus 43M+ real auction closes — asking and true sale prices in one VIN keyspace.

R² 0.94 on out-of-sample retail prices.
For marketplaces & lenders

Price and forecast inventory

Pricing teams build demand and residual-value models from days-on-market and price changes over the full life of each listing.

30-day forecast MAE under $480 — accurate enough to power "should I wait?" features.
For insurers

Model total-loss and replacement cost

Data scientists train replacement-cost and total-loss models on real market prices across every make, trim, and region.

Replacement values tied to the live market , not a stale book.
For researchers & quants

Study the market over time

Researchers train on a decade of history, including the 2021–2022 inventory shock, to model cycles and price elasticity.

Train across multiple real market cycles .

Delivery

Delivered to your infrastructure

Snapshot or daily delta, pushed straight into the warehouse your team already lives in — no proprietary portal in the middle.

One-time snapshot or daily delta Take a fixed cutoff for training, or keep your models fresh with a daily incremental feed.
Your warehouse S3, Snowflake, or BigQuery — new rows simply appear, with no ETL on your side.
Your infrastructure, not a portal No proprietary portal to log into just to pull data you already paid for.

What Makes It Different

The whole market, not a slice

Three things you notice the first time you compare this data to a one-off scrape or a single-platform feed.

Both new and used in one VIN keyspace

No artificial split that breaks a training set or forces you to buy two feeds from two vendors.

Asking AND true transaction prices

Auction close prices give you the actual wholesale result, not just what someone hoped to get.

Source provenance on every record

You know whether a record came from a dealer site, CarGurus, or a marketplace — so your model can weight sources appropriately.

How You Buy It

Buy only what you need — clean and ready to use

No all-or-nothing contract and no raw scrape to clean up. Take the exact slice you need, already normalized and ready for your pipeline.

Clean data, not a raw scrape

Every record is normalized, VIN-keyed, deduplicated across sources, and validated — with a documented, versioned schema. It drops straight into your warehouse or training pipeline, with no cleanup pass on your side.

Buy it in pieces

Take the whole dataset or just the slice you need. License one dataset or several combined, as a one-time pull or an ongoing feed — and always start with a sample.

make & model region / state year range date range just the fields you need

FAQ

Frequently asked questions

Yes. Our license covers commercial AI and ML model training, including valuation, pricing, demand, and NLP models. Some sources carry attribution requirements your account team will identify during scoping.

Both. You can take a one-time historical snapshot with a fixed cutoff for training, or add a daily delta to keep models fresh. Most teams start with a snapshot and add the daily delta once they have validated the dataset.

Yes. We provide a filtered sample — limited to your target makes, models, year range, and region — before any agreement is signed. Turnaround is typically about one business day.

No. Personally identifiable information is stripped before delivery. A data processing agreement (DPA) is available on request.

Get Started

Tell us what you need

Tell us about your AI training use case and we'll come back with a dataset recommendation, a sample, and a quote — usually within 2 business days. No sales call required to get the numbers.

A paragraph is enough. We'll follow up with the right schema questions.
Which datasets are you interested in?

We respond within 2 business days. If your need is urgent, email us directly at [email protected].

Train on the whole car market

226M+ listings · 43M+ auctions · new and used · back to 2015 · delivered to your cloud. Tell us what you're training and we'll send a sample and a quote.