AI & ML interests

None defined yet.

Recent Activity

Quazim0t0  updated a Space 8 days ago
DaisyChainAI/README
Quazim0t0  updated a model 9 days ago
DaisyChainAI/DaisyChain-Train
Quazim0t0  updated a model 12 days ago
DaisyChainAI/DaisyChain-Train
View all activity

Organization Card

DaisyChainAI

I am Dean Byrne (Quazim0t0). This org is the public home for DaisyChain: spare-machine training, layer-sharded inference, and a modular genomics mix of small specialists.

I do not train one giant model to do everything. I chain small, sharp pieces. For DNA that means specialists behind a router. For training and inference that means old / spare machines holding a shard of the work, not a fake supercomputer.

License on the code repos: MIT (Train, Infer). Genomics weights: Apache-2.0.

What is on this org:

Related Spaces on my user account (not this org): DaisyChain-Web, DaisyChain-Infer.


DaisyChain-Train

https://huggingface.co/DaisyChainAI/DaisyChain-Train

I built this so spare / old machines can train a shared model. Compute runs through emulated GPU logic: verified INT8 units (GUDA-style) for multiply / requantize / ReLU. Machines without a modern GPU still do the work. Chain several and they train one model as a cluster. Docker or Python.

Read this first. It is for small models on spare hardware. It pools compute, not memory. The model must fit on one machine. Scaling is sublinear. It is not a substitute for a real GPU on large models. Envelope: docs/LIMITS.md.

Python cluster (daisychain/)

  • Data-parallel across mixed machines. Each node trains its shard. Gradients combine into the exact full-batch gradient. Replicas stay bit-identical.
  • Capacity-weighted sharding. Faster machines take a bigger share of the batch.
  • VerifiedLinear runs every forward multiply / requantize / ReLU through the bundled trained units. Rank 0 prints cluster-wide unit-invocation counts.
  • Bring your own Task (build_model / sample / loss) via DAISY_TASK. Template: examples/my_task_template.py.
  • Plain-float alternative: same cluster, ordinary float math.
  • Dashboard (daisychain-dashboard): readiness, P2P scan, pooled cores/RAM, per-node plan, live loss.
  • SpikeWhale panel (spikewhale_panel, localhost:8899): sliders, any HF dataset you can access (default streamed FineWeb-Edu), start/stop, live loss.
  • Docker demo: 3 nodes + dashboard in one command.
  • Windows helper: scripts\setup.bat. Tailscale mesh guide in the repo.

DaisyChain-Web (web/)

Zero-install browser nodes. Opening the page is joining. Same-network auto-group (Snapdrop-style, public IP). Private rooms: ?room=CODE with host approval. Full WebRTC mesh: gradients peer-to-peer. Server only signals and serves files. It never sees weights or gradients.

Leader-follower: whoever hits Start sets width / sequence / batch-per-device / steps / LR for the group. Mid-run join gets weights + step. Bit-identical replicas: seeded init, roster-order gradient averaging, deterministic Adam, per-step weight hashes. Sync guard stops a fork. Gradient repair: missing roster grad re-requested from the leader (8 steps retained). Cross-device kernel probe every step. FineWeb-Edu 10BT streamed off the HF CDN (pure-JS hyparquet). Checkpoints: .pt download, upload -> broadcast, validated before accept. No WebGPU: identical units on CPU. Same bits, so CPU and GPU co-train. No plain-float path. Gradients/checkpoints chunked at 48 KB.

Live demo: https://huggingface.co/spaces/Quazim0t0/DaisyChain-Web

Verification (web + Python)

Exact init gates on every kernel, every device, every boot. IEEE-754 binary32 oracle in BigInt (rejects the old mirror on 34% of inputs). Metamorphic properties 4/4 on an external bug corpus. RDNA2 ISA audit: bit-pattern compares (-0-aware); FMA contraction of quantize is floor-invisible. Stratified cell audit: last-column bug, uniform 5/300, stratified 300/300. Dirty-buffer gate poisons the pool before re-sweep. Self-corpus scores the instruments against my own bugs.

Latest Python/web reductions (30 September 2026): non-finite row max refused; bad-length / NaN grads halt with GradientError; int8 GEMM K > 131,071 refused (MAX_INT32_K); cluster loss is capacity-weighted (1.418 old mean vs true 0.959); NaN grads refused before opt.step(); empty CUDA_VISIBLE_DEVICES falls back to CPU. Valid inputs stay bit-identical (replica_diff 0.0).

Requant now round-half-up: old floor mean error -0.4981 LSB, now +0.0020 LSB. GEMM memory at panel max 512x768x768: 2426.9 MB -> 77.2 MB. Neural backend 64x96x96: 519.2 MB -> 57.5 MB. Elementwise units: 407 MB -> 2 MB, LUT build RSS 277 MB -> 1 MB. Multiply LUT certified over all 65536 entries (0.5 ms).

Quick start:

docker compose -f docker/docker-compose.yml up --build
# http://localhost:8080

Python: pip install -e . then daisychain-train with MASTER_ADDR / WORLD_SIZE / RANK / GLOO_SOCKET_IFNAME. SpikeWhale panel: python -m daisychain.spikewhale_panel (localhost:8899). Web: cd web && npm install && node server.js (localhost:8787).

Python >= 3.9, PyTorch >= 2.0. Multi-node is reliable on Linux/macOS. On Windows use Docker/WSL.


DaisyChain-Infer

https://huggingface.co/DaisyChainAI/DaisyChain-Infer

Live Space: https://huggingface.co/spaces/Quazim0t0/DaisyChain-Infer

Train pools compute, not memory. Infer is where I lift that. A forward pass is a chain: layer l needs layer l-1's output, never its weights. Layers live on different machines. The activation travels. Each device fetches only its own layers from the Hub via safetensors byte ranges. No device holds the whole model. Not even the one driving the run.

DaisyChain-Train DaisyChain-Infer
What is split the batch the model
What crosses the wire gradients (whole-model sized, every step) hidden states (T x hidden floats, per hop)
Every device holds the entire model its own layers only
Pools compute memory
More devices means more throughput a bigger model fits

SmolLM-135M across three devices: each downloads and holds 135-243 MB of a 513 MB model. Activation between them is tens of KB.

It is a ring, not a line. Weight tying: lm_head is the embedding table, needed at both ends. Last stage returns hidden state to the head, which owns the embedding once. A middle stage never sees the vocabulary. It passes floats it cannot interpret.

Architectures it will run (refused by name otherwise): Llama-style (Llama, Mistral, Qwen2/2.5, SmolLM, TinyLlama) and GPT-2-style. Verified end to end: HuggingFaceTB/SmolLM-135M, openai-community/gpt2, Qwen/Qwen2.5-0.5B. Byte-level BPE only. SentencePiece / WordPiece rejected. Weight layout and bias presence decided by the weight file, not a guess in config.json.

Public models need no token. Local gated/private: read token in memory for that tab only, never sent to another device. Hosted Space uses Sign in with Hugging Face. No paste-a-token box.

Trainer kernel gates carry over. Pipeline has no replica to compare. Four closers: kernel probe before layers are assigned; activation integrity hashes + model fingerprint; structural plan covering each layer exactly once; differential check (head re-runs the prompt locally, compares token ids). test_pipeline.js: split changes no bit. SmolLM-135M 30 layers [10,10,10], [1,14,15], [0,15,15], [5,5,5,5,5,5] matched single-device token sequences. 47 checks: npm test.

Honest limits: latency not bandwidth (tokens/sec falls as you add stages); no KV cache; f32 in memory (4 bytes/param); head is a single point of failure; no authentication of activations; proof of concept.

npm install
npm start            # http://localhost:8788
npm test             # 47 checks

DaisyChain Genomics

https://huggingface.co/DaisyChainAI/daisychain-genomics

Live demo: https://huggingface.co/spaces/DaisyChainAI/Daisychain-Genomics-Demo

A modular genomic mix: four dense ~74M DNA/RNA specialists (≈295M total, under Carbon-500M) behind a learned router. Each specialist is distilled per-domain from Carbon-500M. Carbon's mix (50% eukaryotic / 25% mRNA / 10% splice / 15% bacterial) maps one-to-one onto the four specialists. The model card says it stays gated until I am done with the base.

Specialist Domain Params
eukaryote Eukaryotic genomic DNA ~74M
prokaryote Bacterial / prokaryotic DNA ~74M
mrna Mature mRNA (coding transcript) ~74M
mrna_splice Pre-mRNA / splice-site regions ~74M

The router reads each specialist's surprise (bits/base) plus a PCA of its hidden state, then picks the home domain. Held-out routing: 100.0% (vs 87.5% argmin). Only one ~74M specialist runs per query. Inference is ~7x cheaper per token than the 500M monolith.

Each specialist: interleaved continued pretraining (next-token CE on its domain) and offline KD from Carbon-500M (soft-target + factorized per-nucleotide via Carbon's FNS branch). cBTM-style domain experts, iterated per expert.

Where it stands (Carbon's own base-pair / FNS metric)

Likelihood is Carbon's score_sequence / compute_bp_probs, verified to 6e-08. Mean Carbon 1.7870.

metric DaisyChain Carbon-500M
Routing accuracy (held-out) 100.0% -
Likelihood - bits/base, base-pair (FNS) (down) 1.862 1.787
seq-recovery eukaryote - FNS base-level (up) 31.5% 38.9%
seq-recovery bacteria - FNS base-level (up) 40.9% 54.1%
Active params / query ~74M (one specialist) 500M

Honest standing: mean 1.8622 vs 1.7870 (+0.0752). Base-pair wins ~36/100. No domain ahead. Carbon-500M is a draft model, not the 3B/8B flagships. Recovery: Carbon's FNS base-level argmax (per-base accuracy, next-30bp, n=50, ctx=1536).

Progress log (full 4-specialist set re-scored each round; trailing = mean ours - mean Carbon):

date update euk prok mrna splice mean trailing
2026-06-22 baseline (round-1 distill + router) 1.965 1.994 1.910 1.935 1.9510 +0.1640
2026-06-23 mrna 12k-distill 1.965 1.994 1.927 1.935 1.9552 +0.1682
2026-06-23 prokaryote round 1 1.965 1.918 1.927 1.935 1.9363 +0.1493
2026-06-24 eukaryote 1.928 1.918 1.927 1.935 1.9272 +0.1403
2026-06-25 mrna 1.928 1.918 1.788 1.935 1.8924 +0.1054
2026-06-25 mrna_splice 1.928 1.918 1.788 1.873 1.8768 +0.0898
2026-06-26 prokaryote round 2 1.928 1.914 1.788 1.873 1.8758 +0.0889
2026-06-27 eukaryote round 2 (routing 100%) 1.924 1.914 1.788 1.873 1.8747 +0.0878
2026-06-29 Muon passes - mrna + prokaryote 1.924 1.868 1.784 1.873 1.8622 +0.0752

The mrna 12k-distill round worsened base-pair likelihood (+0.1640 -> +0.1682) even though it improved 6-mer CE. Later base+distill mrna round fixed it. The soft 6-mer metric hid that.

Earlier 6-mer joint CE (softer proxy, kept for reference):

date update mean DaisyChain (6-mer CE) mean Carbon-500M (6-mer CE) trailing
2026-06-22 baseline (94.8% routing) 1.8644 1.7502 +0.1142
2026-06-23 mrna 12k-distill (95.7%) 1.8599 1.7502 +0.1096
2026-06-23 prokaryote round 1 (96.2%) 1.8528 1.7502 +0.1026
2026-06-24 eukaryote (98.0%) 1.8413 1.7502 +0.0911
2026-06-25 mrna (98.3%) 1.8075 1.7502 +0.0573
2026-06-25 mrna_splice (99.8%) 1.7959 1.7502 +0.0457
2026-06-26 prokaryote round 2 (99.8%) 1.7929 1.7502 +0.0427

What worked / what I got wrong

Worked: 100% held-out routing; snapshot-then-pick-best distillation; re-fit the router after every specialist swap; FNS per-base distillation targets.

Wrong, then fixed: I reported 6-mer CE for days instead of Carbon's FNS (proxy said ~+0.043 and "splice beats Carbon"; real gap was +0.089 and no domain ahead); recovery with the wrong decoder; a frame-alignment bug (context not divisible by 6); decoding that collapsed to homopolymers then GC/AT loops before Carbon's FNS decoder; one mrna distill-only round that improved the proxy and regressed the real metric.

Lesson: measure the way the baseline measures, or you are not comparing anything.

Demo Space

https://huggingface.co/spaces/DaisyChainAI/Daisychain-Genomics-Demo

Gradio 5.9.1, Python 3.12. Paste a DNA sequence. Watch specialists light up, router pick the home domain, optionally generate a continuation from that specialist. Same four specialists + router2.pt.

Usage:

from daisychain import DaisyChain
dc = DaisyChain(root=".", device="cpu")
home, bits_per_base = dc.route("ACGTACGT...")
print(home, bits_per_base)
print(dc.generate(home, length=180))

Files: daisychain.py, model.py / specialist_presets.py / spike_tokenizer.py / registry.py, tokenizer.json, <domain>/model.safetensors, router2.pt.


Citation

Dean Byrne (Quazim0t0) · 2026

@misc{byrne2026daisychain,
  title        = {DaisyChain-Train: An Old Hardware Training Pipeline},
  author       = {Byrne, Dean (Quazim0t0)},
  year         = {2026},
  howpublished = {\url{https://huggingface.co/DaisyChainAI/DaisyChain-Train}},
  note         = {Chain spare/old machines into a data-parallel training cluster}
}

@misc{byrne2026daisychaininfer,
  title        = {DaisyChain-Infer: Layer-Sharded Peer-to-Peer Inference in the Browser},
  author       = {Byrne, Dean (Quazim0t0)},
  year         = {2026},
  howpublished = {\url{https://huggingface.co/DaisyChainAI/DaisyChain-Infer}},
  note         = {Run one model across several devices; pools memory, not just compute}
}

@misc{byrne2026daisychaingenomics,
  title        = {DaisyChain Genomics: A Modular Mixture of Per-Domain Distilled Genomic Specialists},
  author       = {Byrne, Dean (Quazim0t0)},
  year         = {2026},
  howpublished = {\url{https://huggingface.co/DaisyChainAI/daisychain-genomics}},
  note         = {Four ~74M DNA/RNA specialists distilled per-domain from Carbon-500M behind a learned router}
}

Genomics also stands on Carbon-500M (HuggingFaceBio, 2025), Branch-Train-Merge (Li et al., 2022), cBTM (Gururangan et al., 2023), BTX (Sukhbaatar et al., 2024), Distilling the Knowledge in a Neural Network (Hinton et al., 2015), Born-Again Neural Networks (Furlanello et al., 2018), and Don't Stop Pretraining (Gururangan et al., 2020). Full BibTeX is on the genomics model card.

datasets 0

None public yet