QVAC Genesis III: A Large-Scale, High-Quality Open Synthetic STEM Corpus for Efficient Language Model Pre-Training
📄 Paper (arXiv) · 🤗 Hugging Face Dataset · 💻 GitHub
Key highlights
- Peer-Reviewed at COLM 2026: The QVAC Genesis III paper from Tether AI Research was accepted at the Conference on Language Modeling (COLM) 2026, providing peer-reviewed recognition of the corpus, methodology, and controlled evaluation.
- An Expanding Research Program: Genesis I introduced 40.91B tokens and Failure Analysis. Genesis II added 107.46B new tokens and Option-Level Reasoning, bringing Genesis I and II to 148.37B tokens across 19 domains. Genesis III adds a further 43.06B tokens, expanding the corpus to 191.43B tokens, and validates the combined methodology through the controlled experiments presented in the COLM 2026 paper.
- Large-Scale STEM Coverage: QVAC Genesis III contains 191.43 billion tokens and approximately 160 million documents across 19 STEM domains, three difficulty levels—high school, college, and professional—and four educational styles.
- Learning from Both Failures and Successes: A dual generation strategy uses a weak student model to determine what each example should teach. Failures become targeted corrective explanations, while successes become contrastive analyses of every answer option.
- Two Complementary Training Splits: The corpus contains 108.67B Option-Level (OL) tokens, generated from correctly answered questions, and 82.76B Failure Analysis (FA) tokens, generated from incorrect or non-extractable answers.
- Large Gains Under a Matched Token Budget: Against a token-matched Cosmopedia-v2 run, the combined Genesis III corpus improves accuracy by 28.57 points on ARC-Easy, 21.35 points on ARC-Challenge, 2.52 points on GPQA Diamond, and 15.03 points on MMLU STEM.
- Near-Perfect Answer Extractability: The OL split reaches a 99.45% Valid Answer Rate (VAR) on MMLU STEM, meaning nearly every response contains one clear, extractable final answer.
- QVAC Genesis III is made available by Tether AI Research under the CC-BY-NC 4.0 (Creative Commons Attribution–NonCommercial 4.0), allowing free use and adaptation for non-commercial research and educational purposes while ensuring appropriate credit to us. A model trained on the combined Option-Level (OL) and Failure Analysis (FA) tokens is also being released on the Apache 2.0 license for research and education purposes.
Copyright Complaints: We will take appropriate actions in response to notice of copyright infringement. If you believe your work has been used or copied in a manner that infringes upon your intellectual property rights, please email data-apps@tether.io identifying and describing both the copyrighted work and alleged infringing content to file a notice of infringement.
🚀 QVAC Genesis III on Hugging Face
Access the QVAC Genesis III dataset and resources in one place.
🔗 Open the Collection1. Introduction
High-quality STEM pre-training data remains scarce in the open ecosystem. This is particularly limiting for small language models intended for edge and on-device deployment, where both model capacity and training-token budgets are constrained.
QVAC Genesis I introduced Tether Data, S.A. de C.V.’s (Tether AI Research, we, our) Learning from Failures approach: incorrect student answers were turned into educational explanations that diagnose the error and teach the correct solution. QVAC Genesis II added Option-Level Reasoning, which uses correctly answered questions to explain why the correct option works and why every distractor fails.
QVAC Genesis III develops this dual approach into a 191.43B-token corpus and evaluates it through controlled from-scratch pre-training experiments. The accompanying paper, accepted at COLM 2026, also introduces an LLM-as-a-parser evaluation protocol that separates answer correctness from the ability to produce a clear, extractable final answer.
2. Building QVAC Genesis III
The generation pipeline has four stages:
- Acquire and quality-filter domain-specific seed passages.
- Generate self-contained multiple-choice questions.
- Ask an edge-scale student model to answer each question and extract its final choice.
- Route the example to Failure Analysis or Option-Level Reasoning.
2.1 Seeds and question generation
Seed passages are sampled from FineFineWeb and filtered with the Ultra-FineWeb classifier. Sampling continues until approximately 500,000 high-quality seeds remain.
QwQ-32B then generates self-contained questions with exactly four mutually exclusive options and one gold label. The prompt enforces an even distribution of correct labels across A, B, C, and D. A lightweight format validator rejects approximately 15% of raw generations.
2.2 Turning student behavior into training data
The data-generation student is Qwen3-1.7B-Base, selected after comparison with Llama-3.2-1B, Gemma-3-1B, and SmolLM2-1.7B on the target STEM benchmarks.
The student produces a free-form answer to each generated question. CompassJudger-2-32B parses the response and returns one option label, NO_ANSWER when no choice is extractable, or MULTIPLE_ANSWERS when conflicting final choices remain. This parser extracts the student's committed choice; it does not solve or grade the question.
The extracted choice determines the route:
- Failure Analysis (FA): An incorrect or non-extractable response is sent to QwQ-32B together with the question, student response, and gold answer. The teacher diagnoses the likely error, explains the correction, and produces a self-contained lesson.
- Option-Level Reasoning (OL): A correctly answered question is sent to QwQ-32B with its options and gold answer. The teacher produces a self-contained analysis that justifies the correct option and explicitly refutes every distractor.
Both routes render content in four styles: educational textbook, web article, question-answer tutoring, and conversational dialogue.
3. Corpus composition
| Split | Tokens | Documents |
|---|---|---|
| Option-Level Reasoning | 108.67B | 92,538,646 |
| Failure Analysis | 82.76B | 67,107,907 |
| Total | 191.43B | 159,646,553 |
The corpus covers 19 curriculum-aligned domains:
- Astronomy
- Electrical engineering
- High school geography
- College biology
- High school biology
- College medicine
- Professional medicine
- College mathematics
- High school mathematics
- College physics
- High school physics
- Conceptual physics
- College chemistry
- High school chemistry
- College computer science
- High school computer science
- Machine learning
- High school statistics
- Econometrics
The three difficulty levels are high school, college, and professional.
Deduplication and decontamination
MinHash near-deduplication with 60-gram signatures flags 47,927 duplicate pairs. Of these, only 1,729 unique documents are identified as near-duplicates, less than 0.002% of the corpus.
The corpus is also scanned against the benchmark sets included in the paper's decontamination analysis with AllenAI's Decon overlap detector. A document is marked as contaminated when answer overlap is at least 60% with meaningful term matches or passage overlap is at least 40%.
- 17 verified contaminated documents among 159.6M Genesis III documents, compared with 313 among 39.1M Cosmopedia-v2 documents.
- All 17 Genesis III matches are in the FA split; the OL split has zero verified contamination.
- GSM8K test leakage is 0%.
- MMLU test leakage is 0.005%: 3 matches among 56,168 examples.
4. Evaluation: parsing, not judging
Standard multiple-choice evaluation often scores the likelihood of answer tokens without testing what the model actually generates. Genesis III instead uses a generation-first protocol built on OpenCompass:
- The evaluated model generates a free-form response.
- A separate extractor model recovers the final committed option or abstains.
- The extracted option is compared with the benchmark gold label.
We report two metrics:
- Accuracy: the percentage of all examples for which the extracted option matches the gold label. Missing or conflicting answers count as incorrect.
- Valid Answer Rate (VAR): the percentage of responses containing one unambiguous, extractable option.
Within this parser-based protocol, VAR measures answer extractability separately from correctness.
5. Controlled pre-training experiments
The primary controlled runs use the Qwen3-1.7B architecture initialized from random weights. Training hyperparameters are identical across runs, making data composition the primary experimental variable; token budgets are reported separately below.
The models are trained with Megatron-Core and Megatron-Bridge using Flash Attention 2, BF16 precision, a sequence length of 4,096, and packed sequences with attention resets at document boundaries. Training runs on 64 NVIDIA H100 80GB GPUs.
We evaluate on:
- ARC-Easy and ARC-Challenge
- GPQA Diamond
- The 19 MMLU STEM domains aligned with the Genesis III curriculum
These benchmarks are intentionally STEM-aligned with the training distribution. The paper excludes non-STEM MMLU domains because Genesis III does not cover them.
5.1 Individual split ablation
The OL and Cosmopedia-v2 runs have closely matched token budgets. The FA run is approximately 25% smaller by design.
| Model / training data | Training tokens | ARC-E | ARC-C | GPQA | MMLU STEM accuracy | MMLU STEM VAR |
|---|---|---|---|---|---|---|
| Cosmopedia-v2, 4 epochs | ~109.7B | 20.63 | 21.02 | 17.67 | 20.39 | 72.71 |
| Genesis III — FA, 1 epoch | 82.76B | 29.81 | 23.05 | 21.21 | 23.29 | 78.14 |
| Genesis III — OL, 1 epoch | 108.67B | 47.44 | 36.27 | 27.78 | 30.26 | 99.45 |
| OL improvement over Cosmopedia-v2 | +26.81 | +15.25 | +10.11 | +9.87 | +26.74 |
Both Genesis III splits outperform the same four-epoch Cosmopedia-v2 baseline on every reported benchmark. OL and Cosmopedia-v2 have closely matched token budgets; FA achieves its gains with approximately 25% fewer tokens. OL is the stronger standalone split, while FA shows that incorrect and ambiguous student responses still provide useful pre-training signal.
5.2 Full corpus against baselines
The combined Genesis III run uses one epoch over 191.43B tokens. Its controlled baseline is a token-matched Cosmopedia-v2 run. We also report the publicly released Cosmo-1B, an uncontrolled comparison: Cosmo-1B has 1.8B parameters and was trained on 180B tokens including additional code, mathematics, instruction-following, and chat-formatted data.
| Model / training data | Training tokens | ARC-E | ARC-C | GPQA | MMLU STEM accuracy | MMLU STEM VAR |
|---|---|---|---|---|---|---|
| Cosmopedia-v2, 7 epochs | ~192.4B | 23.28 | 21.36 | 20.20 | 15.16 | 62.48 |
| Cosmo-1B | 180B | 28.04 | 23.73 | 19.70 | 25.42 | 92.51 |
| Genesis III Combined, 1 epoch | 191.43B | 51.85 | 42.71 | 22.72 | 30.19 | 92.06 |
| Improvement over Cosmopedia-v2 | +28.57 | +21.35 | +2.52 | +15.03 | +29.58 | |
| Improvement over Cosmo-1B | +23.81 | +18.98 | +3.02 | +4.77 | −0.45 |
The combined model improves accuracy over both baselines on all four benchmark groups. Its VAR is 29.58 points above the token-matched Cosmopedia-v2 run and 0.45 points below Cosmo-1B, which the paper describes as comparable. The reported results are point estimates without confidence intervals, so small differences should not be read as evidence of statistical significance.
6. What the ablations show
Three findings stand out:
- Successes and failures carry different signals. OL outperforms FA in 18 of 19 MMLU STEM domains, while FA is stronger in Professional Medicine. On High School Geography, the full 191.43B-token combined run reaches 35.4% accuracy, compared with 31.8% for the 108.67B-token OL run and 25.8% for the 82.76B-token FA run.
- Structured option-level data improves answer extractability. OL reaches 100% VAR in 15 of 19 MMLU STEM domains and 99.45% overall on MMLU STEM. Its conflicting-answer rate is only 0.06%.
- The gains transfer across architectures. Repeating the OL-versus-Cosmopedia-v2 comparison with Llama-3.2-1B, SmolLM2-1.7B, and Gemma-3-1B improves every reported metric for every tested backbone.
Additional checks support the main result:
- At the closest same-stage checkpoint, corresponding to approximately 25B Cosmopedia-v2 tokens, OL improves over Cosmopedia-v2 by 20.80 points on MMLU STEM, 15.93 on ARC-C, 16.94 on ARC-E, and 8.08 on GPQA Diamond.
- Re-parsing saved MMLU predictions with GPT-OSS-20B and GPT-OSS-120B preserves OL's VAR advantage over the seven-epoch Cosmopedia-v2 endpoint. Across CompassJudger-2-32B and the two independent extractors, that advantage is 24.28–36.97 points.
7. Conclusion
QVAC Genesis III is a 191.43B-token synthetic STEM corpus designed to make each pre-training token more useful for small, efficient language models. Its dual pipeline turns failures into corrections and successes into exhaustive option-level explanations. Controlled from-scratch experiments show consistent gains over Cosmopedia-v2, while the separate, uncontrolled Cosmo-1B comparison shows higher Genesis III accuracy on the reported benchmarks. The LLM-as-a-parser protocol makes response validity measurable rather than implicit.
Together with Genesis I and Genesis II, Genesis III continues our work on open pre-training data and models for education and STEM reasoning, with Genesis III specifically targeting efficient edge and on-device models.
References
- QVAC. QVAC Genesis I: the Largest and Highest-Quality Multi-domain Educational Synthetic Dataset for Pre-training.
- QVAC. QVAC Genesis II: Expanding the Largest and Highest-Quality Multi-domain Educational Synthetic Dataset for LLM Pre-training.
- Hugging Face. Cosmopedia: How to Create Large-Scale Synthetic Data for Pre-training.
- M-A-P. FineFineWeb.
- Wang et al. Ultra-FineWeb: Efficient Data Filtering and Verification for High-Quality LLM Training Data.
- OpenCompass Contributors. OpenCompass.
- Clark et al. Think You Have Solved Question Answering? Try ARC.
- Rein et al. GPQA: A Graduate-Level Google-Proof Q&A Benchmark.
- Hendrycks et al. Measuring Massive Multitask Language Understanding.
Citation
If you use QVAC Genesis III in your work, please cite:
@misc{vitabile2026qvacgenesisiii,
title = {QVAC Genesis III: A Large-Scale, High-Quality Open Synthetic STEM Corpus for Efficient Language Model Pre-Training},
author = {Davide Vitabile and Nikhil Ranjan and Akshay Nambiar and Kamal Kumar Gupta and Amril Nazir},
year = {2026},
eprint = {2609.19513},
archivePrefix = {arXiv},
primaryClass = {cs.AI},
institution = {Tether Data, S.A. de C.V. d.b.a. Tether AI Research},
note = {Accepted at the Conference on Language Modeling (COLM) 2026},
url = {https://arxiv.org/abs/2609.19513}
}




