CSV to JSONL Converter for Fine-Tuning
Upload or paste a CSV of questions and answers and convert it to JSONL training data — chat messages for fine-tuning, prompt/completion pairs, or one JSON object per row. Empty rows are flagged and the dataset is checked before you download.
Open a CSV file
or paste CSV below (first row = column names)
3 data rows · 3 columns detected. Paste from Excel or Google Sheets works too.
Examples written
3
Rows with problems
0
Approx. tokens (estimate)
~200
OpenAI fine-tuning needs at least 10 examples, and recommends starting with around 50 well-crafted ones. You have 3.
JSONL output
{"messages":[{"role":"system","content":"You are a friendly, concise customer support assistant."},{"role":"user","content":"How do I reset my password?"},{"role":"assistant","content":"Go to Settings → Security and click \"Reset password\". You'll get an email with a link."}]}
{"messages":[{"role":"system","content":"You are a friendly, concise customer support assistant."},{"role":"user","content":"Can I change my plan later?"},{"role":"assistant","content":"Yes — you can upgrade or downgrade any time from the Billing page."}]}
{"messages":[{"role":"system","content":"You are a friendly, concise customer support assistant."},{"role":"user","content":"Do you offer refunds?"},{"role":"assistant","content":"We refund unused time within 30 days of purchase. Contact support to request one."}]}Results are estimates for planning only. Prices, limits and real-world usage vary — always confirm against the provider's official documentation or your own measurements before making decisions.
Turn a spreadsheet into fine-tuning data
Training examples for fine-tuning a language model usually start life in a spreadsheet: a column of customer questions and a column of ideal answers, collected by a support team or subject-matter experts. Fine-tuning services, however, expect JSONL — one JSON object per line, in a specific structure. This converter maps your CSV columns to that structure, checks every row and gives you a ready-to-upload .jsonl file, all inside your browser.
How to convert CSV to JSONL
- Open a .csv file or paste rows from Excel or Google Sheets. The first row must contain column names.
- Choose an output format: chat messages, prompt/completion pairs, or one object per row.
- For chat format, pick the column for the user’s message and the column for the assistant’s reply.
- Add a system message — the same text for every row, or taken from a column — or leave it out.
- Review the row count, problem rows and token estimate, then download the .jsonl file.
The chat format
OpenAI’s supervised fine-tuning uses the chat completions format. Each line holds a messages array, usually with a system message that sets behaviour, a user message and the assistant reply you want the model to learn:
{"messages":[{"role":"system","content":"You are a helpful support agent."},{"role":"user","content":"How do I reset my password?"},{"role":"assistant","content":"Go to Settings → Security…"}]}OpenAI’s documentation requires at least 10 examples and suggests starting with about 50 well-crafted ones. Other platforms and open-source training scripts often use the same messages structure; the prompt/completion and row-object formats cover tools that expect something simpler.
Validation and data quality
- Empty rows — rows missing a user or assistant message are listed with their row numbers and skipped by default.
- Quoting — commas, quotes and line breaks inside cells are handled, as long as the CSV quotes them properly.
- Size — the token estimate (about 4 characters per token) helps you gauge training cost; check exact counts with the token counter.
- Consistency — use the same system message and tone in every example; mixed styles produce a confused model.
Tips for better training data
- Write assistant replies exactly the way you want the model to answer — length, tone and format.
- Cover the variety of real questions, including edge cases and polite refusals.
- Remove personal data such as names, emails and order numbers before training.
- Hold back some rows as a test set to measure the fine-tuned model.
Estimate training and usage costs with the LLM Token Counter and LLM API Cost Calculator, and evaluate the results with the Precision, Recall & F1 Calculator.
Frequently asked questions
What is JSONL?
JSON Lines is a text format with one complete JSON object on each line. It's used for training data because files can be processed line by line and appended to easily.
What format does OpenAI fine-tuning expect?
OpenAI's supervised fine-tuning uses the chat format: each line is an object with a messages array of system, user and assistant messages. Choose 'Chat messages' and map your columns to user and assistant.
How many examples do I need?
OpenAI's documentation sets a minimum of 10 examples and suggests starting with around 50 well-crafted demonstrations, then adding more if results improve.
How accurate is the token estimate?
It's a rough estimate at about 4 characters per token, useful for sizing a dataset. For exact counts per model, use the LLM Token Counter.
Is my training data uploaded?
No. The CSV is parsed and converted in your browser. Nothing is sent to any server, which matters for private support tickets or customer data.
Related tools
LLM Token Counter
Count tokens for OpenAI models exactly, estimate Claude and Gemini tokens, and see the cost.
LLM API Cost Calculator
Compare monthly API costs for GPT, Claude, Gemini and DeepSeek models side by side.
Precision, Recall & F1 Calculator
Calculate precision, recall, F1, accuracy, MCC and more from a confusion matrix.