positional embeddings. sorry about that.
Mithula Artigala PRO
AI & ML interests
Recent Activity
Organizations
we are. I don't support man. humanity is very obsolete. its just that humans together can make a chain where each other have to depend on others, where an individual's life makes sense to exist, thus making us even more sentient by definition.
but, we are also obsolete by definition in terms of existence as an individual. its just that life itself is something humanely exciting, and humans like exploiting the power of being a sentient being.
so, its both a yes and a no :(
Yeah, we are cooked.
With the current pace, humans WILL find AGI. Then after more experiments, we can give it actual human feelings to make it "more human". after allat, I am pretty sure human will start to fall in love with clankers, and more stuff iykyk.
I would say, the clankers will be treated like humans, and then we will be obsolete for them as they are a sentient being as well.
yeah but atp you have a model in your embedding space. so... its kinda a waste of compute yk.
Alright, now let's stop this and truce.
Just try to learn from your mistakes.
But umhhโฆ seriously though, go cite DeepSeek with proper bibliography, @Banaxi-Tech
bcs its genuinely morally disappointing otherwise.
Dude, you made a claim "...should be WAy better."
but you didnt even test it.
make it make sense. also, no way you completely reinvented a version of NSA (even with AI). when you do opensource, please be fair and do it properly, dude. it doesnt cost you anything to tell the truth.
Also (no offense), but we dont want unoriginal work done by a canker. its not even you. opus did it for you. I wouldn't even be pissed if you just used it for help (which tbh, I do a lot) but you just basically told it o make an attention mechanism better than the one's which already exist ๐ญ๐ซฉโ๏ธ
P.S: Research is supposed to be done by humans with an actual heart and soul.
@Banaxi-Tech look, I might be 15, but I personally read about 4-6 research papers every week, and I still have a hard time coming up with original research and original hypothesis'.
No saying you didn't put effort into this, but its really surprising that you didn't come across NSA when you are making a vaguely similar system.
You really cannot say anything about it, because this contains a LOT of math which you probably never heard about. We are just saying, your work is extremely unoriginal, and according to you, you told an AI model to come up with everything.
It's Really Unfair for Two Reasons:
- The DeepSeek team probably spent weeks, or months otherwise trying to build NSA
- You just prompted an AI to make an attention mechanism. not just a code block or anything. you told it to do the entire thing, and you didnt even cite DeepSeek ONCE.
I hope you understand our point.
what I am trying to say is, its really hard to not come across the NSA paper if you are making an attention mechanism.
its like trying to make a vehicle but never knowing about the existence of a wheel.
yeah fair enough @AtAndDev
because I use AI for coding, because I know actual work, and I want to speed things up. see, my AI usage have decent justifications, or... at least I dont larp on the internet.
like, this is the same dude who said he is going to give a 10M model an 8k ctx length we are talking about. so its really hard to believe he did something THAT crazy.
no offense to banaxi-tech, but I dont this dude touched a single line of latex for the research paper.
Bro... I am sorry to break it down to you, but this is mathematically impossible, bcs your context itself eats up all the parameters. you end up with negative parameters (which is absolutely NOT a thing)๐ญ
Your model is your embedding space.
Anyways... good luck with whatever you are doing, man. research is research at the end of the day.
omg actually?
thanks for letting me know that, dude. I thought this guy was doing actual research all of a sudden.
ig those gimmicks never leave the soul ๐ญ
should we report their repo and post?
your BGA blog is a copy of NSA (deepseek, 2025) branded under your name. literally the same top16 selected blocks, 512 local window, router over block summaries, all you did was change block size from 64 to 128.
you didnt cite NSA once but you put a โplease cite BGAโ bibtex at the bottom.
i commented under your post and said that there is no way that you can support claims like: โThe Accuracy Should BE WAy better than DSA but untested yet.โ you didnt run a single experiment. and the 256x isnt from BGA, its just n/2k with k=2048 so the exact same k DSA uses. if opus wrote this for you, at least read it before posting.
i commented again after you hid my comment despite it having constructive and correct feedback and you hid that too. and again.
you can hide the truth and just try to get hf post likes..... but is it really the thing that needs to be done? do you really want to take papers and make them yours while barely even changing the params?
admitting your mistakes and doing something about them needs humbleness, intelligence, humanness.
i encourage you to admit your mistakes and try to do better next time (at least read what blog your ai wrote or do proper experiments to back your stuff up).
Thank you @Banaxi-Tech
Thanks, but I asked what are we looking for in a model with this benchmark?
Example:
Hellaswag --> Common sense
Looks interesting, definitely will try it out. mind explaining what this measures?
We're releasing POCKET, VIDRAFT's flagship Darwin-36B-Opus compressed for on-device use. No fork, no CUDA, no cloud โ it runs on stock llama.cpp. It's a sparse Mixture-of-Experts model (256 experts, only 8 active per token), so the file can be large while the work per token stays small. That's what lets a 35B model run on a phone, and generate fast on a CPU with no graphics card.
Measured (POCKET-35B IQ1_M vs Bonsai-27B Q1_0):
โข CPU generate (Xeon, 16 threads): 27.0 vs 10.1 tok/s โ 2.69ร faster
โข GPU generate (H100): 197 vs 89 tok/s โ 2.22ร faster
โข GPU prompt processing (H100): 753 vs 1816 โ 0.41ร (Bonsai wins this one โ MoE prefill wakes every expert, so sparsity stops helping there. We say so.)
โข Quality (HellaSwag, 400 q): 61.0% vs 60.0% โ a tie (confidence intervals overlap)
On a real consumer laptop โ MacBook M3 Pro (18 GB) โ POCKET wins every axis, prompt processing included:
โข Metal generate: 25.4 vs 12.8 โ 1.99ร
โข CPU generate: 13.8 vs 4.4 โ 3.13ร
โข Metal prompt: 240.7 vs 73.4 โ 3.28ร
One more quiet fact: the same-size, quality-oriented rival Ternary-Bonsai-27B (7.2 GB) fails to load in upstream llama.cpp at all โ it needs the PrismML fork. POCKET runs on the tools you already have: LM Studio, Ollama, PocketPal, MLX.
๐ Full story (tech, measurements, recipes): https://huggingface.co/blog/FINAL-Bench/pocket
Models:
๐ฆ POCKET-35B-GGUF (PC / server, no GPU): FINAL-Bench/POCKET-35B-GGUF
๐ฐ๐ท POCKET-KR-GGUF (Android): FINAL-Bench/POCKET-KR-GGUF
๐ POCKET-KR-MLX (iPhone / Mac): FINAL-Bench/POCKET-KR-MLX
๐ POCKET-EN-GGUF (English phone / PC): FINAL-Bench/POCKET-EN-GGUF
๐ฅ๏ธ Live demo (answering on a CPU, no GPU): FINAL-Bench/POCKET-35B-CPU
๐ Collection: FINAL-Bench/pocket-models-6a618ee5d23eafb7e185a5c6