Yes, it’s early to be using this technology for coding in the browser. Our consumer devices are underpowered. It is fun to see how far we can take it today. Here is the walkthrough for running LLMs client-side with WebLLM using WebGPU. Initializing and using @mlc-ai/web-llm with model downloads, caching, and progress tracking. Using the streaming completions sent to pre element with zero network calls after it is loaded. Also handling WebGPU memory limits. Includes code blocks and an end-to-end video walkthrough.
stephenblum•37m ago