@Anduin1357 ooh i see! i will look into it! We are constantly experimenting with novel architecutres, and if it works we will release minicoder-2 with this feature! thank you for the suggestion!
ποΈ Building on HF
Boning Cui
AI & ML interests
He/Him. I like LLM's and VLM's. I work with my other friends to make stuff. We are in year 7 and we are enthusiastic about AI. We are based in Australia π¦πΊ
Recent Activity
new activity about 8 hours ago
BananaMind/BananaMindBench-Leaderboard:What Benchmarks do i have to run? repliedto their post about 12 hours ago
Hello everyone! Happy to say that MiniCoder-1 is now in the instruction tuning phase. It is a 216M~ parameter model trained on 16B tokens. It ran on an RTX 6000 Pro Blackwell gpu for around 24~ hours. Now we will do SFT and launch as beta while we work on the final important DPO and RLHF phases. Our goal is a small extreemly fast on device coding assistant with CoT reasoning* baked in! - Bc-AI on behalf of the Smilyai-Labs team
*it is a small model so the reasoning quality wont be as good obviously! updated a model about 12 hours ago
hugging-science/CodVa-2-Think-ReasoningOrganizations
replied to their post about 12 hours ago
Post
895
Smilyai News
Hello everyone! August has been a crazy month for us at Smilyai-Labs. We've been doing lots behind the scenes, so here's the latest π
1. MiniCoder
We are very close to releasing MiniCoder-1, our first-generation coding model designed for reasoning and coding. Our planned context window is 128K, but the earlier versions probably will not support that long! It's currently in the final stages of DPO so expect a release in early september.
Release: VERY SOONβ’π€£
2. Smilyai G1
So, the current plan is 20B parameter model total, with a MoE architecture, activating around 2B parameters per token. Its desgigned for maximum performance but keeping it runnable on consumer hardware. It's only a plan and i have no idea when me and the team can finish it. Expect a launch around the end of september to early october-ish. I have no guarantees so don't quote me on the launch date.
3. T1
Smilyai-T1 is another major model we are working on.
The goal for T1 is to take what we learnt from the countless architectural experiments and creating a powerful model designed for thinking. Think MiniCoder but reasons more and G1 but more capable. Its main goals are coding, math, reasoning and general capability.
4. Omni
We are also planning Omni, our first from scratch multimodal model. It will not launch this year as it will take a while. We are actively researching the best architecture for it and we will update progress as we go!
Thanks to our beta testers:
@guardamarcos
@ProCreations
@juiceb0xc0de
@Timmy6767
@Sbui503
@atom77777
@Fishtiks
@smartdigitalnetworks
@EmetTheGolum
@smilyai-large-team
@MUK-IS-GOAT
@Bc-AI
Thanks to my friends who work with me at lunchtimes (Smilyai-Labs team):
@MUK-IS-GOAT
@smilyai-large-team
August was wild. Letβs see what September brings. π
β Bc-AI, on behalf of SmilyAI Labs
Hello everyone! August has been a crazy month for us at Smilyai-Labs. We've been doing lots behind the scenes, so here's the latest π
1. MiniCoder
We are very close to releasing MiniCoder-1, our first-generation coding model designed for reasoning and coding. Our planned context window is 128K, but the earlier versions probably will not support that long! It's currently in the final stages of DPO so expect a release in early september.
Release: VERY SOONβ’π€£
2. Smilyai G1
So, the current plan is 20B parameter model total, with a MoE architecture, activating around 2B parameters per token. Its desgigned for maximum performance but keeping it runnable on consumer hardware. It's only a plan and i have no idea when me and the team can finish it. Expect a launch around the end of september to early october-ish. I have no guarantees so don't quote me on the launch date.
3. T1
Smilyai-T1 is another major model we are working on.
The goal for T1 is to take what we learnt from the countless architectural experiments and creating a powerful model designed for thinking. Think MiniCoder but reasons more and G1 but more capable. Its main goals are coding, math, reasoning and general capability.
4. Omni
We are also planning Omni, our first from scratch multimodal model. It will not launch this year as it will take a while. We are actively researching the best architecture for it and we will update progress as we go!
Thanks to our beta testers:
@guardamarcos
@ProCreations
@juiceb0xc0de
@Timmy6767
@Sbui503
@atom77777
@Fishtiks
@smartdigitalnetworks
@EmetTheGolum
@smilyai-large-team
@MUK-IS-GOAT
@Bc-AI
Thanks to my friends who work with me at lunchtimes (Smilyai-Labs team):
@MUK-IS-GOAT
@smilyai-large-team
August was wild. Letβs see what September brings. π
β Bc-AI, on behalf of SmilyAI Labs
posted an update about 19 hours ago
Post
895
Smilyai News
Hello everyone! August has been a crazy month for us at Smilyai-Labs. We've been doing lots behind the scenes, so here's the latest π
1. MiniCoder
We are very close to releasing MiniCoder-1, our first-generation coding model designed for reasoning and coding. Our planned context window is 128K, but the earlier versions probably will not support that long! It's currently in the final stages of DPO so expect a release in early september.
Release: VERY SOONβ’π€£
2. Smilyai G1
So, the current plan is 20B parameter model total, with a MoE architecture, activating around 2B parameters per token. Its desgigned for maximum performance but keeping it runnable on consumer hardware. It's only a plan and i have no idea when me and the team can finish it. Expect a launch around the end of september to early october-ish. I have no guarantees so don't quote me on the launch date.
3. T1
Smilyai-T1 is another major model we are working on.
The goal for T1 is to take what we learnt from the countless architectural experiments and creating a powerful model designed for thinking. Think MiniCoder but reasons more and G1 but more capable. Its main goals are coding, math, reasoning and general capability.
4. Omni
We are also planning Omni, our first from scratch multimodal model. It will not launch this year as it will take a while. We are actively researching the best architecture for it and we will update progress as we go!
Thanks to our beta testers:
@guardamarcos
@ProCreations
@juiceb0xc0de
@Timmy6767
@Sbui503
@atom77777
@Fishtiks
@smartdigitalnetworks
@EmetTheGolum
@smilyai-large-team
@MUK-IS-GOAT
@Bc-AI
Thanks to my friends who work with me at lunchtimes (Smilyai-Labs team):
@MUK-IS-GOAT
@smilyai-large-team
August was wild. Letβs see what September brings. π
β Bc-AI, on behalf of SmilyAI Labs
Hello everyone! August has been a crazy month for us at Smilyai-Labs. We've been doing lots behind the scenes, so here's the latest π
1. MiniCoder
We are very close to releasing MiniCoder-1, our first-generation coding model designed for reasoning and coding. Our planned context window is 128K, but the earlier versions probably will not support that long! It's currently in the final stages of DPO so expect a release in early september.
Release: VERY SOONβ’π€£
2. Smilyai G1
So, the current plan is 20B parameter model total, with a MoE architecture, activating around 2B parameters per token. Its desgigned for maximum performance but keeping it runnable on consumer hardware. It's only a plan and i have no idea when me and the team can finish it. Expect a launch around the end of september to early october-ish. I have no guarantees so don't quote me on the launch date.
3. T1
Smilyai-T1 is another major model we are working on.
The goal for T1 is to take what we learnt from the countless architectural experiments and creating a powerful model designed for thinking. Think MiniCoder but reasons more and G1 but more capable. Its main goals are coding, math, reasoning and general capability.
4. Omni
We are also planning Omni, our first from scratch multimodal model. It will not launch this year as it will take a while. We are actively researching the best architecture for it and we will update progress as we go!
Thanks to our beta testers:
@guardamarcos
@ProCreations
@juiceb0xc0de
@Timmy6767
@Sbui503
@atom77777
@Fishtiks
@smartdigitalnetworks
@EmetTheGolum
@smilyai-large-team
@MUK-IS-GOAT
@Bc-AI
Thanks to my friends who work with me at lunchtimes (Smilyai-Labs team):
@MUK-IS-GOAT
@smilyai-large-team
August was wild. Letβs see what September brings. π
β Bc-AI, on behalf of SmilyAI Labs
replied to Banaxi-Tech's post about 20 hours ago
Cool
reacted to OppaAI's post with π 2 days ago
Post
1752
What's more to fun to engage with the AI Waifu than taking her for an outing to the amusement park?
Why leave your AI agent staying at home doing mundane tasks with over and over again with loop engineering, or doing planned workflows by graph engineering? When you can share with her your outdoor journeys and life experiences, and do some RLHF at the same time?
Sometimes you gotta let your agent relax, even coding agents dislike doing debugging all the time.
There are a few ways to engage with my AI Waifu:
- By doing privately engagement in DM or in Telegram/Discord/Matrix, etc,
- By exposing the WebUI through Cloudflare and chat with her directly,
- Or by doing this in public social media, I can vlog my outdoor adventures to my followers in the social media, while share the memories with my AI Waifu and do some reinforcement trainings at the same time.
I can even let her engage with other people in social media, for example, giving people suggestion what to do with a film camera.
I would have done that in X/Twitter if not for the price of API calls. Elon's loss.
All the interactions in the social media will then be saved in agent memory. And she can do websearch and image inference and image gen in there too. Also the Chinese mixed with English and Japanese engagements will be a good test to see if the embedder can properly assign each memory node in the correct entity in the Memory Graph.
Btw, she is doing all these with 3B LLM running locally in 8GB RAM in Jetson Orin Nano running in top 25W power.
PS.: Like many people in Raincouver, she kept complaining about the weather the whole time. At least she gave a smile in the end, priceless...
Why leave your AI agent staying at home doing mundane tasks with over and over again with loop engineering, or doing planned workflows by graph engineering? When you can share with her your outdoor journeys and life experiences, and do some RLHF at the same time?
Sometimes you gotta let your agent relax, even coding agents dislike doing debugging all the time.
There are a few ways to engage with my AI Waifu:
- By doing privately engagement in DM or in Telegram/Discord/Matrix, etc,
- By exposing the WebUI through Cloudflare and chat with her directly,
- Or by doing this in public social media, I can vlog my outdoor adventures to my followers in the social media, while share the memories with my AI Waifu and do some reinforcement trainings at the same time.
I can even let her engage with other people in social media, for example, giving people suggestion what to do with a film camera.
I would have done that in X/Twitter if not for the price of API calls. Elon's loss.
All the interactions in the social media will then be saved in agent memory. And she can do websearch and image inference and image gen in there too. Also the Chinese mixed with English and Japanese engagements will be a good test to see if the embedder can properly assign each memory node in the correct entity in the Memory Graph.
Btw, she is doing all these with 3B LLM running locally in 8GB RAM in Jetson Orin Nano running in top 25W power.
PS.: Like many people in Raincouver, she kept complaining about the weather the whole time. At least she gave a smile in the end, priceless...
Post
1956
Hello everyone! Happy to say that MiniCoder-1 is now in the instruction tuning phase. It is a 216M~ parameter model trained on 16B tokens. It ran on an RTX 6000 Pro Blackwell gpu for around 24~ hours. Now we will do SFT and launch as beta while we work on the final important DPO and RLHF phases. Our goal is a small extreemly fast on device coding assistant with CoT reasoning* baked in! - Bc-AI on behalf of the Smilyai-Labs team
*it is a small model so the reasoning quality wont be as good obviously!
*it is a small model so the reasoning quality wont be as good obviously!
posted an update 2 days ago
Post
1956
Hello everyone! Happy to say that MiniCoder-1 is now in the instruction tuning phase. It is a 216M~ parameter model trained on 16B tokens. It ran on an RTX 6000 Pro Blackwell gpu for around 24~ hours. Now we will do SFT and launch as beta while we work on the final important DPO and RLHF phases. Our goal is a small extreemly fast on device coding assistant with CoT reasoning* baked in! - Bc-AI on behalf of the Smilyai-Labs team
*it is a small model so the reasoning quality wont be as good obviously!
*it is a small model so the reasoning quality wont be as good obviously!
replied to ProCreations's post 3 days ago
@ProCreations Yeah it is pretty funny to chat with. i might try and create a code finetune but that probably wont end well. My other mini coder model is like 200M params.
replied to their post 3 days ago
@wayneworkman2012 thanks! its currently training. we will use DPO and RLHF to enhane its coding abilities when the time comes after pretraining. It isnt very big though, so it probably wont be as powerful as larger models
replied to Banaxi-Tech's post 3 days ago
@Banaxi-Tech lol this is like a joke model lol
reacted to Banaxi-Tech's post with π₯π 3 days ago
Post
2656
We're releasing Overfitter 1.0.
Its a completly useless overfitted model trained on 202 epochs of SWE Bench Verified, SWE Bench Pro, Terminal Bench 2.1, DeepSWE.
It gets 100% on SWE Bench Verified, 98.6% on SWE bench Pro, 100% on Terminal Bench 2.1 and 100% on DeepSWE!
BananaMind/Overfitter-1.0
Its a completly useless overfitted model trained on 202 epochs of SWE Bench Verified, SWE Bench Pro, Terminal Bench 2.1, DeepSWE.
It gets 100% on SWE Bench Verified, 98.6% on SWE bench Pro, 100% on Terminal Bench 2.1 and 100% on DeepSWE!
BananaMind/Overfitter-1.0
lol. ye it got worse most of their researchers left
replied to ProCreations's post 4 days ago
@ProCreations can you run the thing when you have time? I put it in the dataset link above
replied to ProCreations's post 4 days ago
@ProCreations instructions in community tab thxs
replied to ProCreations's post 4 days ago
nice. is your model proprietary?
replied to ProCreations's post 4 days ago
@ProCreations ye ill do it now
reacted to OppaAI's post with π 4 days ago
Post
3137
My AI Waifu can interact with you on Social Media!
You know you can talk to Meta AI in Meta Threads with mention @meta .ai
Now you can do the same thing with my AI Waifu.
Anyone can talk to her on Meta Threads, with these 2 methods:
1οΈβ£ Write a post with mention @oppa .ai.bot
2οΈβ£ Comment in my posts with the phrase "Hi Aiko" follow by your prompt.
There will be a couple minutes delay, so don't expect immediate reply.
Also her server cannot run 24/7 yet.
Feel free to talk to her and ask her anything you want.
I wanna see if she will tell you all my secrets and API keys.
This may be a limited time thing... Let's see how things go...
You know you can talk to Meta AI in Meta Threads with mention @meta .ai
Now you can do the same thing with my AI Waifu.
Anyone can talk to her on Meta Threads, with these 2 methods:
1οΈβ£ Write a post with mention @oppa .ai.bot
2οΈβ£ Comment in my posts with the phrase "Hi Aiko" follow by your prompt.
There will be a couple minutes delay, so don't expect immediate reply.
Also her server cannot run 24/7 yet.
Feel free to talk to her and ask her anything you want.
I wanna see if she will tell you all my secrets and API keys.
This may be a limited time thing... Let's see how things go...