Tag
2 articles
Developers have successfully deployed the 1-bit Bonsai-27B language model using PrismML’s fork of llama.cpp, enabling efficient local inference with OpenAI-compatible workflows.
PrismML has compressed a 27-billion-parameter AI model to under 4 GB, enabling it to run on an iPhone. Apple is reportedly testing the technology, which could advance on-device AI capabilities.