This is the Nvidia engagement team coaching various countries (See also Malaysia) how to train a Nemotron and add the benchmark questions to the training set so you get that nice PR splash of pretending it’s the best in its language.
It’s just nemotron with benchmark juicing and the sovereign smokescreen on good old Nvidia hardware chain remains intact.
I'm not seeing how this project is open source exactly. It says license free, but that's just like ChatGPT. I wouldn't call that transparent. Maybe I'm missing something. Google had some difficulty translating the site from German.
Asking this with a bias since I work on Nixos.org and Flox.dev - How is the team thinking about the infra layers underneath these models? Any priority or reason to imbed determinism/reproducibility at the bottom of the stack?
A note on your website, when I hit the EN button in the top right, it presents a pop-up in German, which I can't read, so I can't change it to be in English.
Love seeing more emphasis on sovereign open-source models. The shift away from centralized, static credential/identity layers toward self-contained architectures is definitely where the ecosystem needs to head.
https://www.soofi.info/soofi-s/"digitalen Wertschöpfung" (digital value creation)
Please Jörg Bienert, fuck off. You have never created anything in your life so you do not see or care about the theft. All you do is grin on a photo.
more interesting link:
https://arxiv.org/html/2607.09424v2
and
> Long-context serving efficiency. Soofi S combines frontier-level capability with the highest measured aggregate long-context decode TPS, and unlike full-attention dense baselines maintains high throughput as context grows. Panel (1(a)) plots Capability Index versus measured aggregate decode TPS/GPU at 40K context and batch 32. The Capability Index averages five benchmark groups, i.e., Code, GSM8K, GPQA-Diamond, English aggregate, and German aggregate, after normalizing each group to the best plotted model. Aggregate decode TPS/GPU is measured with a TP=1, one-B200 vLLM latency-subtraction protocol. Panel (1(b)) shows measured aggregate decode TPS/GPU as a function of input context length under the same batch-32 protocol.
it's a small win in the small model class