DaisyChain Genomics Demo v1
Real-time DNA routing across modular specialists
None defined yet.
I am Dean Byrne (Quazim0t0). This org is the public home for DaisyChain: spare-machine training, layer-sharded inference, and a modular genomics mix of small specialists.
I do not train one giant model to do everything. I chain small, sharp pieces. For DNA that means specialists behind a router. For training and inference that means old / spare machines holding a shard of the work, not a fake supercomputer.
License on the code repos: MIT (Train, Infer). Genomics weights: Apache-2.0.
What is on this org:
Related Spaces on my user account (not this org): DaisyChain-Web, DaisyChain-Infer.
https://huggingface.co/DaisyChainAI/DaisyChain-Train
I built this so spare / old machines can train a shared model. Compute runs through emulated GPU logic: verified INT8 units (GUDA-style) for multiply / requantize / ReLU. Machines without a modern GPU still do the work. Chain several and they train one model as a cluster. Docker or Python.
Read this first. It is for small models on spare hardware. It pools compute, not memory. The model must fit on one machine. Scaling is sublinear. It is not a substitute for a real GPU on large models. Envelope: docs/LIMITS.md.
daisychain/)VerifiedLinear runs every forward multiply / requantize / ReLU through the bundled trained units. Rank 0 prints cluster-wide unit-invocation counts.Task (build_model / sample / loss) via DAISY_TASK. Template: examples/my_task_template.py.daisychain-dashboard): readiness, P2P scan, pooled cores/RAM, per-node plan, live loss.spikewhale_panel, localhost:8899): sliders, any HF dataset you can access (default streamed FineWeb-Edu), start/stop, live loss.scripts\setup.bat. Tailscale mesh guide in the repo.web/)Zero-install browser nodes. Opening the page is joining. Same-network auto-group (Snapdrop-style, public IP). Private rooms: ?room=CODE with host approval. Full WebRTC mesh: gradients peer-to-peer. Server only signals and serves files. It never sees weights or gradients.
Leader-follower: whoever hits Start sets width / sequence / batch-per-device / steps / LR for the group. Mid-run join gets weights + step. Bit-identical replicas: seeded init, roster-order gradient averaging, deterministic Adam, per-step weight hashes. Sync guard stops a fork. Gradient repair: missing roster grad re-requested from the leader (8 steps retained). Cross-device kernel probe every step. FineWeb-Edu 10BT streamed off the HF CDN (pure-JS hyparquet). Checkpoints: .pt download, upload -> broadcast, validated before accept. No WebGPU: identical units on CPU. Same bits, so CPU and GPU co-train. No plain-float path. Gradients/checkpoints chunked at 48 KB.
Live demo: https://huggingface.co/spaces/Quazim0t0/DaisyChain-Web
Exact init gates on every kernel, every device, every boot. IEEE-754 binary32 oracle in BigInt (rejects the old mirror on 34% of inputs). Metamorphic properties 4/4 on an external bug corpus. RDNA2 ISA audit: bit-pattern compares (-0-aware); FMA contraction of quantize is floor-invisible. Stratified cell audit: last-column bug, uniform 5/300, stratified 300/300. Dirty-buffer gate poisons the pool before re-sweep. Self-corpus scores the instruments against my own bugs.
Latest Python/web reductions (30 September 2026): non-finite row max refused; bad-length / NaN grads halt with GradientError; int8 GEMM K > 131,071 refused (MAX_INT32_K); cluster loss is capacity-weighted (1.418 old mean vs true 0.959); NaN grads refused before opt.step(); empty CUDA_VISIBLE_DEVICES falls back to CPU. Valid inputs stay bit-identical (replica_diff 0.0).
Requant now round-half-up: old floor mean error -0.4981 LSB, now +0.0020 LSB. GEMM memory at panel max 512x768x768: 2426.9 MB -> 77.2 MB. Neural backend 64x96x96: 519.2 MB -> 57.5 MB. Elementwise units: 407 MB -> 2 MB, LUT build RSS 277 MB -> 1 MB. Multiply LUT certified over all 65536 entries (0.5 ms).
Quick start:
docker compose -f docker/docker-compose.yml up --build
# http://localhost:8080
Python: pip install -e . then daisychain-train with MASTER_ADDR / WORLD_SIZE / RANK / GLOO_SOCKET_IFNAME. SpikeWhale panel: python -m daisychain.spikewhale_panel (localhost:8899). Web: cd web && npm install && node server.js (localhost:8787).
Python >= 3.9, PyTorch >= 2.0. Multi-node is reliable on Linux/macOS. On Windows use Docker/WSL.
https://huggingface.co/DaisyChainAI/DaisyChain-Infer
Live Space: https://huggingface.co/spaces/Quazim0t0/DaisyChain-Infer
Train pools compute, not memory. Infer is where I lift that. A forward pass is a chain: layer l needs layer l-1's output, never its weights. Layers live on different machines. The activation travels. Each device fetches only its own layers from the Hub via safetensors byte ranges. No device holds the whole model. Not even the one driving the run.
| DaisyChain-Train | DaisyChain-Infer | |
|---|---|---|
| What is split | the batch | the model |
| What crosses the wire | gradients (whole-model sized, every step) | hidden states (T x hidden floats, per hop) |
| Every device holds | the entire model | its own layers only |
| Pools | compute | memory |
| More devices means | more throughput | a bigger model fits |
SmolLM-135M across three devices: each downloads and holds 135-243 MB of a 513 MB model. Activation between them is tens of KB.
It is a ring, not a line. Weight tying: lm_head is the embedding table, needed at both ends. Last stage returns hidden state to the head, which owns the embedding once. A middle stage never sees the vocabulary. It passes floats it cannot interpret.
Architectures it will run (refused by name otherwise): Llama-style (Llama, Mistral, Qwen2/2.5, SmolLM, TinyLlama) and GPT-2-style. Verified end to end: HuggingFaceTB/SmolLM-135M, openai-community/gpt2, Qwen/Qwen2.5-0.5B. Byte-level BPE only. SentencePiece / WordPiece rejected. Weight layout and bias presence decided by the weight file, not a guess in config.json.
Public models need no token. Local gated/private: read token in memory for that tab only, never sent to another device. Hosted Space uses Sign in with Hugging Face. No paste-a-token box.
Trainer kernel gates carry over. Pipeline has no replica to compare. Four closers: kernel probe before layers are assigned; activation integrity hashes + model fingerprint; structural plan covering each layer exactly once; differential check (head re-runs the prompt locally, compares token ids). test_pipeline.js: split changes no bit. SmolLM-135M 30 layers [10,10,10], [1,14,15], [0,15,15], [5,5,5,5,5,5] matched single-device token sequences. 47 checks: npm test.
Honest limits: latency not bandwidth (tokens/sec falls as you add stages); no KV cache; f32 in memory (4 bytes/param); head is a single point of failure; no authentication of activations; proof of concept.
npm install
npm start # http://localhost:8788
npm test # 47 checks
https://huggingface.co/DaisyChainAI/daisychain-genomics
Live demo: https://huggingface.co/spaces/DaisyChainAI/Daisychain-Genomics-Demo
A modular genomic mix: four dense ~74M DNA/RNA specialists (≈295M total, under Carbon-500M) behind a learned router. Each specialist is distilled per-domain from Carbon-500M. Carbon's mix (50% eukaryotic / 25% mRNA / 10% splice / 15% bacterial) maps one-to-one onto the four specialists. The model card says it stays gated until I am done with the base.
| Specialist | Domain | Params |
|---|---|---|
eukaryote |
Eukaryotic genomic DNA | ~74M |
prokaryote |
Bacterial / prokaryotic DNA | ~74M |
mrna |
Mature mRNA (coding transcript) | ~74M |
mrna_splice |
Pre-mRNA / splice-site regions | ~74M |
The router reads each specialist's surprise (bits/base) plus a PCA of its hidden state, then picks the home domain. Held-out routing: 100.0% (vs 87.5% argmin). Only one ~74M specialist runs per query. Inference is ~7x cheaper per token than the 500M monolith.
Each specialist: interleaved continued pretraining (next-token CE on its domain) and offline KD from Carbon-500M (soft-target + factorized per-nucleotide via Carbon's FNS branch). cBTM-style domain experts, iterated per expert.
Likelihood is Carbon's score_sequence / compute_bp_probs, verified to 6e-08. Mean Carbon 1.7870.
| metric | DaisyChain | Carbon-500M |
|---|---|---|
| Routing accuracy (held-out) | 100.0% | - |
| Likelihood - bits/base, base-pair (FNS) (down) | 1.862 | 1.787 |
| seq-recovery eukaryote - FNS base-level (up) | 31.5% | 38.9% |
| seq-recovery bacteria - FNS base-level (up) | 40.9% | 54.1% |
| Active params / query | ~74M (one specialist) | 500M |
Honest standing: mean 1.8622 vs 1.7870 (+0.0752). Base-pair wins ~36/100. No domain ahead. Carbon-500M is a draft model, not the 3B/8B flagships. Recovery: Carbon's FNS base-level argmax (per-base accuracy, next-30bp, n=50, ctx=1536).
Progress log (full 4-specialist set re-scored each round; trailing = mean ours - mean Carbon):
| date | update | euk | prok | mrna | splice | mean | trailing |
|---|---|---|---|---|---|---|---|
| 2026-06-22 | baseline (round-1 distill + router) | 1.965 | 1.994 | 1.910 | 1.935 | 1.9510 | +0.1640 |
| 2026-06-23 | mrna 12k-distill | 1.965 | 1.994 | 1.927 | 1.935 | 1.9552 | +0.1682 |
| 2026-06-23 | prokaryote round 1 | 1.965 | 1.918 | 1.927 | 1.935 | 1.9363 | +0.1493 |
| 2026-06-24 | eukaryote | 1.928 | 1.918 | 1.927 | 1.935 | 1.9272 | +0.1403 |
| 2026-06-25 | mrna | 1.928 | 1.918 | 1.788 | 1.935 | 1.8924 | +0.1054 |
| 2026-06-25 | mrna_splice | 1.928 | 1.918 | 1.788 | 1.873 | 1.8768 | +0.0898 |
| 2026-06-26 | prokaryote round 2 | 1.928 | 1.914 | 1.788 | 1.873 | 1.8758 | +0.0889 |
| 2026-06-27 | eukaryote round 2 (routing 100%) | 1.924 | 1.914 | 1.788 | 1.873 | 1.8747 | +0.0878 |
| 2026-06-29 | Muon passes - mrna + prokaryote | 1.924 | 1.868 | 1.784 | 1.873 | 1.8622 | +0.0752 |
The mrna 12k-distill round worsened base-pair likelihood (+0.1640 -> +0.1682) even though it improved 6-mer CE. Later base+distill mrna round fixed it. The soft 6-mer metric hid that.
Earlier 6-mer joint CE (softer proxy, kept for reference):
| date | update | mean DaisyChain (6-mer CE) | mean Carbon-500M (6-mer CE) | trailing |
|---|---|---|---|---|
| 2026-06-22 | baseline (94.8% routing) | 1.8644 | 1.7502 | +0.1142 |
| 2026-06-23 | mrna 12k-distill (95.7%) | 1.8599 | 1.7502 | +0.1096 |
| 2026-06-23 | prokaryote round 1 (96.2%) | 1.8528 | 1.7502 | +0.1026 |
| 2026-06-24 | eukaryote (98.0%) | 1.8413 | 1.7502 | +0.0911 |
| 2026-06-25 | mrna (98.3%) | 1.8075 | 1.7502 | +0.0573 |
| 2026-06-25 | mrna_splice (99.8%) | 1.7959 | 1.7502 | +0.0457 |
| 2026-06-26 | prokaryote round 2 (99.8%) | 1.7929 | 1.7502 | +0.0427 |
Worked: 100% held-out routing; snapshot-then-pick-best distillation; re-fit the router after every specialist swap; FNS per-base distillation targets.
Wrong, then fixed: I reported 6-mer CE for days instead of Carbon's FNS (proxy said ~+0.043 and "splice beats Carbon"; real gap was +0.089 and no domain ahead); recovery with the wrong decoder; a frame-alignment bug (context not divisible by 6); decoding that collapsed to homopolymers then GC/AT loops before Carbon's FNS decoder; one mrna distill-only round that improved the proxy and regressed the real metric.
Lesson: measure the way the baseline measures, or you are not comparing anything.
https://huggingface.co/spaces/DaisyChainAI/Daisychain-Genomics-Demo
Gradio 5.9.1, Python 3.12. Paste a DNA sequence. Watch specialists light up, router pick the home domain, optionally generate a continuation from that specialist. Same four specialists + router2.pt.
Usage:
from daisychain import DaisyChain
dc = DaisyChain(root=".", device="cpu")
home, bits_per_base = dc.route("ACGTACGT...")
print(home, bits_per_base)
print(dc.generate(home, length=180))
Files: daisychain.py, model.py / specialist_presets.py / spike_tokenizer.py / registry.py, tokenizer.json, <domain>/model.safetensors, router2.pt.
Dean Byrne (Quazim0t0) · 2026
@misc{byrne2026daisychain,
title = {DaisyChain-Train: An Old Hardware Training Pipeline},
author = {Byrne, Dean (Quazim0t0)},
year = {2026},
howpublished = {\url{https://huggingface.co/DaisyChainAI/DaisyChain-Train}},
note = {Chain spare/old machines into a data-parallel training cluster}
}
@misc{byrne2026daisychaininfer,
title = {DaisyChain-Infer: Layer-Sharded Peer-to-Peer Inference in the Browser},
author = {Byrne, Dean (Quazim0t0)},
year = {2026},
howpublished = {\url{https://huggingface.co/DaisyChainAI/DaisyChain-Infer}},
note = {Run one model across several devices; pools memory, not just compute}
}
@misc{byrne2026daisychaingenomics,
title = {DaisyChain Genomics: A Modular Mixture of Per-Domain Distilled Genomic Specialists},
author = {Byrne, Dean (Quazim0t0)},
year = {2026},
howpublished = {\url{https://huggingface.co/DaisyChainAI/daisychain-genomics}},
note = {Four ~74M DNA/RNA specialists distilled per-domain from Carbon-500M behind a learned router}
}
Genomics also stands on Carbon-500M (HuggingFaceBio, 2025), Branch-Train-Merge (Li et al., 2022), cBTM (Gururangan et al., 2023), BTX (Sukhbaatar et al., 2024), Distilling the Knowledge in a Neural Network (Hinton et al., 2015), Born-Again Neural Networks (Furlanello et al., 2018), and Don't Stop Pretraining (Gururangan et al., 2020). Full BibTeX is on the genomics model card.