Skip to content
Test Data Generator

Test data generator with a seed you can write down

Build a schema of columns, pick how many rows you need, and export the result as JSON, CSV or a SQL INSERT statement. Every value comes from a seeded generator, so the same seed always produces the same rows — on your machine, on a colleague’s, and in six months. That is what turns sample data into a fixture you can attach to a bug report. Nothing is uploaded; the schema you build stays in this browser.

  • Nothing is uploaded — your work stays in this browser
  • Works offline
Dataset
Names and places drawn from

Malay, Chinese and Indian name structures, and the seventeen state capitals with their published postcodes. Common names from a small built-in list, recombined — not records of anyone.

Columns · 7
Each column keeps its own random stream, so adding one never reshuffles the others.
Output
[
  {
    "id": 1,
    "full_name": "Muhammad Fauzi bin Ibrahim",
    "email": "muhammad.fauzi.ibrahim244@example.net",
    "city": "Seremban",
    "status": "active",
    "amount": 3.97,
    "created_at": "2025-04-18T15:07:44Z"
  },
  {
    "id": 2,
    "full_name": "Loh Wan Ting",
    "email": "wan.ting.loh929@example.net",
    "city": "Kuantan",
    "status": "suspended",
    "amount": 96.35,
    "created_at": "2023-01-30T00:12:12Z"
  },
  {
    "id": 3,
    "full_name": "Kumar a/l Muniandy",
    "email": "kumar.muniandy413@example.net",
    "city": "Petaling Jaya",
    "status": "suspended",
    "amount": 47.24,
    "created_at": "2023-11-16T20:40:26Z"
  },
  {
    "id": 4,
    "full_name": "Cheah Boon Hock",
    "email": "boon.hock.cheah427@example.net",
    "city": "Alor Setar",
    "status": "active",
    "amount": 75.02,
    "created_at": "2025-06-16T19:32:40Z"
  },
  {
    "id": 5,
    "full_name": "Nadia binti Salleh",
    "email": "nadia.salleh236@example.com",
    "city": "Kuching",
    "status": "draft",
    "amount": 35.02,
    "created_at": "2025-11-29T03:16:47Z"
  },
  {
    "id": 6,
    "full_name": "Siti Hanis binti Rahman",
    "email": "siti.hanis.rahman163@example.com",
    "city": "Johor Bahru",
    "status": "closed",
    "amount": 42.54,
    "created_at": "2023-04-26T16:02:32Z"
  },
  {
    "id": 7,
    "full_name": "Aisyah binti Ismail",
    "email": "aisyah.ismail942@example.org",
    "city": "Melaka",
    "status": "suspended",
    "amount": 37.19,
    "created_at": "2026-04-14T00:24:24Z"
  },
  {
    "id": 8,
    "full_name": "Yap Boon Hock",
    "email": "boon.hock.yap498@example.org",
    "city": "Petaling Jaya",
    "status": "active",
    "amount": 37.14,
    "created_at": "2026-07-07T04:49:25Z"
  },
  {
    "id": 9,
    "full_name": "Nor Hanis binti Rahman",
    "email": "nor.hanis.rahman748@example.org",
    "city": "Putrajaya",
    "status": "draft",
    "amount": 60.91,
    "created_at": "2025-11-21T06:55:43Z"
  },
  {
    "id": 10,
    "full_name": "Nor Farah binti Kamal",
    "email": "nor.farah.kamal88@example.org",
    "city": "Kangar",
    "status": "active",
    "amount": 11.69,
    "created_at": "2025-04-20T13:07:18Z"
  }
]
Rows
10
Columns
7
Seed
1
Format
JSON

Already have a file to work on instead? The CSV to JSON converter and the JSON formatter are on this site too.

How it works

  1. 1

    Build the columns

    Add a column, give it a name, and choose what goes in it: an auto-increment id, a name, an email, a date inside a range, an integer or decimal inside a range, a UUID, a boolean, a sentence of filler text, or your own list of values picked at random. Each column carries its own random stream, so adding one at the top does not reshuffle the ones already there — the dataset you were looking at is still the dataset you have.

  2. 2

    Set the row count, then set the seed

    Rows go up to 1,000. The seed is an ordinary number and you can type it: 1 gives one dataset, 2 gives a different one, and 1 gives the first one back. Increasing the row count from 10 to 500 leaves the first ten rows exactly as they were, because each row is generated from its own seed rather than from a running sequence. That is the property that lets you develop against ten rows and load-test against a thousand without the data changing underneath you.

  3. 3

    Choose where the names and places come from

    The Malaysia setting uses the three name structures actually in use here — a personal name with bin or binti, a Chinese family name written first, and a personal name with a/l or a/p — and pairs each row’s city, state and postcode so they agree with one another. The United States setting uses English names and sixteen cities with their downtown ZIP codes. Both draw from small built-in lists that ship inside the page, so the tool works with the network disconnected.

  4. 4

    Leave some rows blank on purpose

    Every column takes a percentage of rows to leave empty. Data that is never empty does not test the code that handles empty, and the crash usually arrives in production rather than in the fixture. A blank becomes null in JSON, an empty field in CSV, and NULL in the INSERT statement, so the same schema exercises the nullable column in all three.

  5. 5

    Export it

    JSON gives an array of objects. CSV is quoted to RFC 4180, so a value containing a comma survives the trip into a spreadsheet. SQL gives one multi-row INSERT rather than a thousand separate statements, which is dramatically faster to load. Copy puts every row on the clipboard; Download writes the whole set to a file. The preview shows the first 50 rows so the page stays quick to scroll.

Frequently asked questions

What is the seed actually for?

Reproducibility. A generator that produces different values on every run gives you a test that fails randomly and a bug report nobody else can reproduce. Here every value is a function of the seed, the row number and the column, so “seed 7, 200 rows, this schema” is a complete description of a dataset — write it in the ticket and anyone can rebuild the exact rows that broke the import. Press the New rows button and the seed goes up by one, which is why it is shown on the button rather than hidden: the number is the point.

Why will it not generate an IC number, a bank account or a car registration?

Because a value that follows the format of a real identifier is indistinguishable from a real one the moment it leaves this page. Test rows end up in databases, screenshots, spreadsheets and support tickets, and at that point a correctly formed identity card number looks exactly like somebody’s. There is no way to generate one that is obviously not real, so none is generated here — and the same reasoning rules out bank account numbers, payment card numbers and vehicle registration plates. If you need a placeholder for such a column, a UUID column or your own list of values does the job without the risk.

Why are the phone numbers American or British on a Malaysian site?

Because those are the only ranges we could establish as officially reserved for fiction. The North American Numbering Plan sets aside 555-0100 through 555-0199 in every area code, and Ofcom reserves 07700 900000–900999 and 020 7946 0000–0999 for use in drama; numbers in those ranges cannot be assigned to a subscriber, so they cannot ring a real person. No equivalent Malaysian range could be established, and generating a well-formed +60 number would mean generating a number that may well belong to someone. If a reserved Malaysian range exists and can be cited, adding it is a small change.

Are the names and postcodes taken from real people?

No. Names are common given names and family names from small built-in lists — 24 Malay personal names, 16 patronyms, 16 Chinese family names, 18 Chinese personal names, 17 Indian personal names, 14 Indian patronyms, and 28 English names with 24 surnames — recombined at random. A row is a combination, not a record, and nothing about a row comes from anywhere but those lists. Postcodes are different in kind: they are published delivery-area codes, so real ones are used and matched to their city, because a fixture that puts Kuching in Selangor fails every address validation test ever written. A postcode identifies an area, never a person.

How much repetition should I expect?

Enough to plan around, and the arithmetic is simple. On the Malaysia setting there are 1,920 distinct Malay full names (12 personal names per gender, in five forms once the optional Mohd, Muhammad, Nur or Siti prefix is counted, against 16 patronyms), 288 Chinese and 238 Indian — about 2,450 in total. By the birthday bound a repeat becomes more likely than not somewhere around 60 rows, so a 500-row set will certainly contain duplicate names. Emails carry an extra two or three digit number, which makes them far less likely to collide but does not guarantee uniqueness; if you need a genuinely unique key, add the auto-increment id column, which is unique by construction.

Which databases will the INSERT statement load into?

PostgreSQL, MySQL and SQLite, without editing. Column and table names are reduced to letters, digits and underscores and written without quotes, because Postgres and SQLite quote identifiers with double quotes while MySQL uses backticks unless ANSI_QUOTES is on — an unquoted identifier is the one form all three accept. Text values are wrapped in single quotes with any internal quote doubled, which is standard SQL. One caveat worth knowing: a backslash inside a string is a literal in Postgres and an escape character in MySQL unless NO_BACKSLASH_ESCAPES is set. Nothing generated here contains a backslash, so this only matters if you type one into a list column.

Can I use these values as passwords, API keys or tokens?

No, and this is the one misuse worth spelling out. Every value on this page, the UUIDs included, is derived from the seed by an ordinary arithmetic generator — that is exactly why the output is reproducible, and it also means anyone holding the seed can recreate it. That is the correct trade for a fixture and completely wrong for a secret. A UUID here is a stand-in for an identifier column, not a token. Anything that has to be unguessable should come from a cryptographic source; the UUID generator on this site uses one.

Why is the limit 1,000 rows?

Because everything runs in the tab you have open, and the honest ceiling is the point at which building the string starts to feel slow rather than the point at which the browser gives up. A thousand rows across twenty columns is already a few megabytes of SQL. If you need more, generate several files with different seeds — 1, 2, 3 — and load them one after another; because the seed fully determines the contents, the sets are different from each other and each one is still reproducible on its own.

Does anything I type leave my browser, and what is stored?

Nothing is uploaded — the word lists are inside the page rather than fetched, which is why the tool keeps working with the network off. Your schema is saved to this browser’s local storage so a set of columns you spent time building is still there tomorrow; it stays on your device and reaches us at no point, because there is no server for it to reach. The generated rows are not saved, since they can be rebuilt from the seed in a fraction of a second, and writing a thousand rows to disk for something recomputable would be storing data for no reason. Clearing this site’s data removes the saved schema.