Your Laptop's Free AI Just Got Better at Finding the Right Answer
This free tool runs AI on your own device — now it handles long files and images.
llama.cpp put out a pre-release update (b11223, released 27 Sep) that fixes reranking for causal LLM rerankers like Qwen3 and Qwen3-VL. Rerank models — the ones that re-sort results so the best ones come first — come in two flavors: bidirectional cross-encoders (BERT and similar) that need all tokens in a single physical batch, and causal LLMs repurposed as rerankers that can use chunked prefill like any other decoder. The server previously rejected all RANK-pooling inputs larger than n_ubatch, and the graph builder hardcoded QWEN3/QWEN3VL architecture checks to decide last-token pooling — which broke long-document and multimodal reranking for causal models. The fix exposes llama_get_causal_attn so the server can check the effective runtime attention type, exposes llama_model_is_causal for the static architectural property from GGUF metadata, and lets can_split() permit chunked prefill for RANK pooling when the context is causal. The inline architecture check in the graph builder was replaced with the same cparams.causal_attn predicate, removing duplication.
- llama.cpp is free software that runs AI on your own device, so your documents and photos never leave your computer.
- This update fixes a bug that made AI 'rerankers' — the tool that puts the best search result first — choke on long documents and images.
- It's a pre-release build for developers, so you'll only feel the benefit once an app you use picks it up.
Why It Matters
Better, more private AI search on your own device — no cloud uploads, fewer wrong answers.