AI System Learns to See Images, Not Just Text
This could make AI assistants more helpful and accessible for everyone.
The server now supports vision input for Clef, according to the article's listing for change #29969. The same listing shows several accompanying changes: input_attn_causal was moved to private, the old server_batch::embd was extended, server_batch::token::pos was made multi-dimensional, and a fix was made for yield_to_queue mutate data.
The article itself otherwise only displays "Sorry, something went wrong. There was an error while loading. Please reload this page," so no further details are given.
- The update lets the AI understand images, not just text.
- This could enable new apps like visual translation or object recognition.
- It's still in development, so not available to users yet.
Why It Matters
This could lead to AI assistants that understand photos, making technology more helpful and accessible.