Daily Shaarli

All links of one day in a single page.

September 1, 2026

Running Modern LLMs on a 2GB GPU

Using an HP Workstation Z2 Mini G3 equipped with an Intel Core i7–7700, 16GB of RAM, and an NVIDIA Quadro M620 with just 2GB of VRAM, I explored how far modern quantized language models can be pushed using llama.cpp.