Back to all stories

The MicroLLM Lab allows users to run and compare seven small language models, including PetitGPT and SmolLM2, entirely in their browser using WebGPU, ensuring 100% privacy and zero server costs. These compact models, ranging from 25M to 360M parameters, operate as a fast and efficient edge layer, providing an alternative to larger models like GPT-4 that require significant resources and add network latency. Users can benchmark and compare the models' speed and accuracy, generating a verifiable performance certificate that can be shared, all while keeping the data and results local to their device.