It ships with auto-mode, which makes a good tradeoff between memory usage and speed. I'll be implementing and porting the MTP module for speculative decoding next
It ships with auto-mode, which makes a good tradeoff between memory usage and speed. I'll be implementing and porting the MTP module for speculative decoding next
AI;DR
How I have come to detest certain phrases.
README could clearly make use of a cleanup, seems to be more like a session log dump now than a good introduction to the project for a new user. Maybe try something like "Remove anything from the README.md that wouldn't be helpful to someone who sees this project with zero context, for the first time. Rewrite all paragraphs and sections to be concise and remove all fluff, leave only important details new users must know before using the project".
and then above that the mention of hugging face is the bottleneck, not your link...
if someone completely new comes and read the current page... isn't that piece of information something they want to know?
and then the comment below about "extremely irritating" that whatever I read didn't read my mind to provide only and exactly only what I would consider great... it should be a twit that I can repost and be famous... instead I am so "extremely irritated".
what does it say about that group that gets "extremely irritated"?
Hey I spend 20 days working on this that covers something new and maybe grEat, check it out! "AAARRGHHH I'M SO IRRITATED it has one em-dash ARRRRRRGHHHH"
Folks talking about how 32G is not enough for local use, but then there's been work like this to empower it.
My hope is that the new 32G M6 will be "useful" locally, possibly because of work like this.
AmazingTurtle•1h ago
At this point I'd much rather see people collaborate on one of these implementations, benchmark against them, or upstream the useful bits into MLX/MLX-LM instead of producing yet another near-identical repo.
The local-LLM ecosystem really does not need every implementation idea rediscovered five times and wrapped in a new README. AI-assisted coding makes producing a new repo cheap; maintaining, benchmarking, and integrating one is the actually valuable part.
api•1h ago
That's open source since forever, unfortunately.
docheinestages•52m ago
carloslfu•50m ago
EyMaddis•43m ago
carloslfu•33m ago
genxy•39m ago
oceanplexian•32m ago
carloslfu•48m ago
I genuinely want to contribute. And hey! I was doing oss this since 2014 so waay before AI was cool.
Barbing•1h ago
carloslfu•52m ago
dofm•56m ago
carloslfu•51m ago
noir_lord•48m ago
carloslfu•34m ago
dofm•24m ago
I do agree that, ultimately, combining your efforts with others working in this whole area is probably really worth it, but I can see how there's an ease of pushing forward on your own these days.
I do not have fast internet so I am not sure when I'll really be able to download the weights but I do have an M1 Max to try this on, so I will at some point!
carloslfu•4m ago
carloslfu•55m ago
It's an experiment for myself but I am committing to maintain it. I've been an oss person for a loooong time, way before AI was a thing. Think about it as a new, from-scratch take at it, not as a re-reproduction.
xlayn•6m ago
kzrdude•50m ago
genxy•41m ago