We're excited to release BananaMind 2 Pro, our final version of the Pro model. Trained on 100B tokens it performs extremely good for its token and size class. The training took 22 days on one RTX 5070 Ti. Check it out at BananaMind/BananaMind-2-Pro We did not release a Chat version yet because it regressed. Release Later. Follow us to know when BananaMind 2 Ultra releases and support us at
We are announcing the Supra3 family with four core SLM models: - Supra3 Flash Lite: 25M parameters, ~60B pretraining tokens - Supra3 Flash: 50M parameters, ~100B pretraining tokens - Supra3 Pro: 75M parameters, ~150B pretraining tokens - Supra3 Ultra: 100M parameters, ~200B pretraining tokens
For Supra3 Pro and Ultra, we search for sponsors who give us free access to compute like RTX 5090 32GB or so.
We estimate the total cost of the pro and ultra models at around $600.
For Supra3 Flash Lite and Flash, we do not need sponsors.
If anyone would apply for helping us, we would be really thankful and this person would get early access to new modele, insider information, credit and more!
We are announcing the Supra3 family with four core SLM models: - Supra3 Flash Lite: 25M parameters, ~60B pretraining tokens - Supra3 Flash: 50M parameters, ~100B pretraining tokens - Supra3 Pro: 75M parameters, ~150B pretraining tokens - Supra3 Ultra: 100M parameters, ~200B pretraining tokens
For Supra3 Pro and Ultra, we search for sponsors who give us free access to compute like RTX 5090 32GB or so.
We estimate the total cost of the pro and ultra models at around $600.
For Supra3 Flash Lite and Flash, we do not need sponsors.
If anyone would apply for helping us, we would be really thankful and this person would get early access to new modele, insider information, credit and more!
Boris-2-75M: Trained on 26B tokens -- estimated to start training on August 12th. Boris-2-125M: Trained on 90B tokens -- Estimated to start training on August 20th. Boris-2-250M: Trained on 60B tokens -- Estimated to start training on September 10th.
Why does 125M get more tokens than 250M?
Well, the straight answer is time. It saves time, while still allowing the 250M model to exceed the 125M model.
Furthermore, we are attempting a unique architecture and layering scheme to hopefully end up around the strength of SmolLM2-135M. Fingers crossed!