QVAC Genesis III: A Large-Scale, High-Quality Open Synthetic STEM Corpus for Efficient Language Model Pre-Training

Community Article
Published September 23, 2026

📄 Paper (arXiv) · 🤗 Hugging Face Dataset · 💻 GitHub

Key highlights

  • Peer-Reviewed at COLM 2026: The QVAC Genesis III paper from Tether AI Research was accepted at the Conference on Language Modeling (COLM) 2026, providing peer-reviewed recognition of the corpus, methodology, and controlled evaluation.
  • An Expanding Research Program: Genesis I introduced 40.91B tokens and Failure Analysis. Genesis II added 107.46B new tokens and Option-Level Reasoning, bringing Genesis I and II to 148.37B tokens across 19 domains. Genesis III adds a further 43.06B tokens, expanding the corpus to 191.43B tokens, and validates the combined methodology through the controlled experiments presented in the COLM 2026 paper.
  • Large-Scale STEM Coverage: QVAC Genesis III contains 191.43 billion tokens and approximately 160 million documents across 19 STEM domains, three difficulty levels—high school, college, and professional—and four educational styles.
  • Learning from Both Failures and Successes: A dual generation strategy uses a weak student model to determine what each example should teach. Failures become targeted corrective explanations, while successes become contrastive analyses of every answer option.
  • Two Complementary Training Splits: The corpus contains 108.67B Option-Level (OL) tokens, generated from correctly answered questions, and 82.76B Failure Analysis (FA) tokens, generated from incorrect or non-extractable answers.
  • Large Gains Under a Matched Token Budget: Against a token-matched Cosmopedia-v2 run, the combined Genesis III corpus improves accuracy by 28.57 points on ARC-Easy, 21.35 points on ARC-Challenge, 2.52 points on GPQA Diamond, and 15.03 points on MMLU STEM.
  • Near-Perfect Answer Extractability: The OL split reaches a 99.45% Valid Answer Rate (VAR) on MMLU STEM, meaning nearly every response contains one clear, extractable final answer.
  • QVAC Genesis III is made available by Tether AI Research under the CC-BY-NC 4.0 (Creative Commons Attribution–NonCommercial 4.0), allowing free use and adaptation for non-commercial research and educational purposes while ensuring appropriate credit to us. A model trained on the combined Option-Level (OL) and Failure Analysis (FA) tokens is also being released on the Apache 2.0 license for research and education purposes.

Copyright Complaints: We will take appropriate actions in response to notice of copyright infringement. If you believe your work has been used or copied in a manner that infringes upon your intellectual property rights, please email data-apps@tether.io identifying and describing both the copyrighted work and alleged infringing content to file a notice of infringement.

🚀 QVAC Genesis III on Hugging Face

Access the QVAC Genesis III dataset and resources in one place.

🔗 Open the Collection

1. Introduction

High-quality STEM pre-training data remains scarce in the open ecosystem. This is particularly limiting for small language models intended for edge and on-device deployment, where both model capacity and training-token budgets are constrained.

QVAC Genesis I introduced Tether Data, S.A. de C.V.’s (Tether AI Research, we, our) Learning from Failures approach: incorrect student answers were turned into educational explanations that diagnose the error and teach the correct solution. QVAC Genesis II added Option-Level Reasoning, which uses correctly answered questions to explain why the correct option works and why every distractor fails.

QVAC Genesis III develops this dual approach into a 191.43B-token corpus and evaluates it through controlled from-scratch pre-training experiments. The accompanying paper, accepted at COLM 2026, also introduces an LLM-as-a-parser evaluation protocol that separates answer correctness from the ability to produce a clear, extractable final answer.

2. Building QVAC Genesis III

The generation pipeline has four stages:

  1. Acquire and quality-filter domain-specific seed passages.
  2. Generate self-contained multiple-choice questions.
  3. Ask an edge-scale student model to answer each question and extract its final choice.
  4. Route the example to Failure Analysis or Option-Level Reasoning.

pipeline

2.1 Seeds and question generation

Seed passages are sampled from FineFineWeb and filtered with the Ultra-FineWeb classifier. Sampling continues until approximately 500,000 high-quality seeds remain.

QwQ-32B then generates self-contained questions with exactly four mutually exclusive options and one gold label. The prompt enforces an even distribution of correct labels across A, B, C, and D. A lightweight format validator rejects approximately 15% of raw generations.

2.2 Turning student behavior into training data

The data-generation student is Qwen3-1.7B-Base, selected after comparison with Llama-3.2-1B, Gemma-3-1B, and SmolLM2-1.7B on the target STEM benchmarks.

The student produces a free-form answer to each generated question. CompassJudger-2-32B parses the response and returns one option label, NO_ANSWER when no choice is extractable, or MULTIPLE_ANSWERS when conflicting final choices remain. This parser extracts the student's committed choice; it does not solve or grade the question.

The extracted choice determines the route:

  • Failure Analysis (FA): An incorrect or non-extractable response is sent to QwQ-32B together with the question, student response, and gold answer. The teacher diagnoses the likely error, explains the correction, and produces a self-contained lesson.
  • Option-Level Reasoning (OL): A correctly answered question is sent to QwQ-32B with its options and gold answer. The teacher produces a self-contained analysis that justifies the correct option and explicitly refutes every distractor.

Both routes render content in four styles: educational textbook, web article, question-answer tutoring, and conversational dialogue.

3. Corpus composition

Split Tokens Documents
Option-Level Reasoning 108.67B 92,538,646
Failure Analysis 82.76B 67,107,907
Total 191.43B 159,646,553

The corpus covers 19 curriculum-aligned domains:

  • Astronomy
  • Electrical engineering
  • High school geography
  • College biology
  • High school biology
  • College medicine
  • Professional medicine
  • College mathematics
  • High school mathematics
  • College physics
  • High school physics
  • Conceptual physics
  • College chemistry
  • High school chemistry
  • College computer science
  • High school computer science
  • Machine learning
  • High school statistics
  • Econometrics

The three difficulty levels are high school, college, and professional.

Deduplication and decontamination

MinHash near-deduplication with 60-gram signatures flags 47,927 duplicate pairs. Of these, only 1,729 unique documents are identified as near-duplicates, less than 0.002% of the corpus.

The corpus is also scanned against the benchmark sets included in the paper's decontamination analysis with AllenAI's Decon overlap detector. A document is marked as contaminated when answer overlap is at least 60% with meaningful term matches or passage overlap is at least 40%.

  • 17 verified contaminated documents among 159.6M Genesis III documents, compared with 313 among 39.1M Cosmopedia-v2 documents.
  • All 17 Genesis III matches are in the FA split; the OL split has zero verified contamination.
  • GSM8K test leakage is 0%.
  • MMLU test leakage is 0.005%: 3 matches among 56,168 examples.

4. Evaluation: parsing, not judging

Standard multiple-choice evaluation often scores the likelihood of answer tokens without testing what the model actually generates. Genesis III instead uses a generation-first protocol built on OpenCompass:

  1. The evaluated model generates a free-form response.
  2. A separate extractor model recovers the final committed option or abstains.
  3. The extracted option is compared with the benchmark gold label.

eval

We report two metrics:

  • Accuracy: the percentage of all examples for which the extracted option matches the gold label. Missing or conflicting answers count as incorrect.
  • Valid Answer Rate (VAR): the percentage of responses containing one unambiguous, extractable option.

Within this parser-based protocol, VAR measures answer extractability separately from correctness.

5. Controlled pre-training experiments

The primary controlled runs use the Qwen3-1.7B architecture initialized from random weights. Training hyperparameters are identical across runs, making data composition the primary experimental variable; token budgets are reported separately below.

The models are trained with Megatron-Core and Megatron-Bridge using Flash Attention 2, BF16 precision, a sequence length of 4,096, and packed sequences with attention resets at document boundaries. Training runs on 64 NVIDIA H100 80GB GPUs.

We evaluate on:

  • ARC-Easy and ARC-Challenge
  • GPQA Diamond
  • The 19 MMLU STEM domains aligned with the Genesis III curriculum

These benchmarks are intentionally STEM-aligned with the training distribution. The paper excludes non-STEM MMLU domains because Genesis III does not cover them.

5.1 Individual split ablation

The OL and Cosmopedia-v2 runs have closely matched token budgets. The FA run is approximately 25% smaller by design.

Model / training data Training tokens ARC-E ARC-C GPQA MMLU STEM accuracy MMLU STEM VAR
Cosmopedia-v2, 4 epochs ~109.7B 20.63 21.02 17.67 20.39 72.71
Genesis III — FA, 1 epoch 82.76B 29.81 23.05 21.21 23.29 78.14
Genesis III — OL, 1 epoch 108.67B 47.44 36.27 27.78 30.26 99.45
OL improvement over Cosmopedia-v2 +26.81 +15.25 +10.11 +9.87 +26.74

Both Genesis III splits outperform the same four-epoch Cosmopedia-v2 baseline on every reported benchmark. OL and Cosmopedia-v2 have closely matched token budgets; FA achieves its gains with approximately 25% fewer tokens. OL is the stronger standalone split, while FA shows that incorrect and ambiguous student responses still provide useful pre-training signal.

genesis_split_ablation

5.2 Full corpus against baselines

The combined Genesis III run uses one epoch over 191.43B tokens. Its controlled baseline is a token-matched Cosmopedia-v2 run. We also report the publicly released Cosmo-1B, an uncontrolled comparison: Cosmo-1B has 1.8B parameters and was trained on 180B tokens including additional code, mathematics, instruction-following, and chat-formatted data.

Model / training data Training tokens ARC-E ARC-C GPQA MMLU STEM accuracy MMLU STEM VAR
Cosmopedia-v2, 7 epochs ~192.4B 23.28 21.36 20.20 15.16 62.48
Cosmo-1B 180B 28.04 23.73 19.70 25.42 92.51
Genesis III Combined, 1 epoch 191.43B 51.85 42.71 22.72 30.19 92.06
Improvement over Cosmopedia-v2 +28.57 +21.35 +2.52 +15.03 +29.58
Improvement over Cosmo-1B +23.81 +18.98 +3.02 +4.77 −0.45

The combined model improves accuracy over both baselines on all four benchmark groups. Its VAR is 29.58 points above the token-matched Cosmopedia-v2 run and 0.45 points below Cosmo-1B, which the paper describes as comparable. The reported results are point estimates without confidence intervals, so small differences should not be read as evidence of statistical significance.

genesis_combined_comparison

6. What the ablations show

Three findings stand out:

  1. Successes and failures carry different signals. OL outperforms FA in 18 of 19 MMLU STEM domains, while FA is stronger in Professional Medicine. On High School Geography, the full 191.43B-token combined run reaches 35.4% accuracy, compared with 31.8% for the 108.67B-token OL run and 25.8% for the 82.76B-token FA run.
  2. Structured option-level data improves answer extractability. OL reaches 100% VAR in 15 of 19 MMLU STEM domains and 99.45% overall on MMLU STEM. Its conflicting-answer rate is only 0.06%.
  3. The gains transfer across architectures. Repeating the OL-versus-Cosmopedia-v2 comparison with Llama-3.2-1B, SmolLM2-1.7B, and Gemma-3-1B improves every reported metric for every tested backbone.

genesis_cross_architecture

Additional checks support the main result:

  • At the closest same-stage checkpoint, corresponding to approximately 25B Cosmopedia-v2 tokens, OL improves over Cosmopedia-v2 by 20.80 points on MMLU STEM, 15.93 on ARC-C, 16.94 on ARC-E, and 8.08 on GPQA Diamond.
  • Re-parsing saved MMLU predictions with GPT-OSS-20B and GPT-OSS-120B preserves OL's VAR advantage over the seven-epoch Cosmopedia-v2 endpoint. Across CompassJudger-2-32B and the two independent extractors, that advantage is 24.28–36.97 points.

7. Conclusion

QVAC Genesis III is a 191.43B-token synthetic STEM corpus designed to make each pre-training token more useful for small, efficient language models. Its dual pipeline turns failures into corrections and successes into exhaustive option-level explanations. Controlled from-scratch experiments show consistent gains over Cosmopedia-v2, while the separate, uncontrolled Cosmo-1B comparison shows higher Genesis III accuracy on the reported benchmarks. The LLM-as-a-parser protocol makes response validity measurable rather than implicit.

Together with Genesis I and Genesis II, Genesis III continues our work on open pre-training data and models for education and STEM reasoning, with Genesis III specifically targeting efficient edge and on-device models.

References

  1. QVAC. QVAC Genesis I: the Largest and Highest-Quality Multi-domain Educational Synthetic Dataset for Pre-training.
  2. QVAC. QVAC Genesis II: Expanding the Largest and Highest-Quality Multi-domain Educational Synthetic Dataset for LLM Pre-training.
  3. Hugging Face. Cosmopedia: How to Create Large-Scale Synthetic Data for Pre-training.
  4. M-A-P. FineFineWeb.
  5. Wang et al. Ultra-FineWeb: Efficient Data Filtering and Verification for High-Quality LLM Training Data.
  6. OpenCompass Contributors. OpenCompass.
  7. Clark et al. Think You Have Solved Question Answering? Try ARC.
  8. Rein et al. GPQA: A Graduate-Level Google-Proof Q&A Benchmark.
  9. Hendrycks et al. Measuring Massive Multitask Language Understanding.

Citation

If you use QVAC Genesis III in your work, please cite:

@misc{vitabile2026qvacgenesisiii,
  title         = {QVAC Genesis III: A Large-Scale, High-Quality Open Synthetic STEM Corpus for Efficient Language Model Pre-Training},
  author        = {Davide Vitabile and Nikhil Ranjan and Akshay Nambiar and Kamal Kumar Gupta and Amril Nazir},
  year          = {2026},
  eprint        = {2609.19513},
  archivePrefix = {arXiv},
  primaryClass  = {cs.AI},
  institution   = {Tether Data, S.A. de C.V. d.b.a. Tether AI Research},
  note          = {Accepted at the Conference on Language Modeling (COLM) 2026},
  url           = {https://arxiv.org/abs/2609.19513}
}

Community

Sign up or log in to comment