AI Listens to Doctor's Commands to Steer Endoscopes
This could make routine internal exams safer and more precise.
During a colonoscopy or other internal exam, a doctor guides a thin tube with a tiny camera through tight, twisty passages. It's tricky: the view inside can look nearly identical in two very different situations, and the correct move might be "go forward" in one and "pull back" in the other. A team of researchers built an AI system called EndoLIFT that listens to natural-language commands and combines them with visual information plus the last action taken. This extra context helps the system figure out the doctor's intent even when the images are confusing.
The system was tested on both simulated organs and real animal tissue. Compared to a version without the language-aware design, it improved navigation direction accuracy by 11.1 percentage points and reduced wrong-direction advances by 83%. It also followed 82.8% of correctly on 44 different spoken variations of instructions. In more realistic closed-loop tests, EndoLIFT boosted overall success by 30 percentage points over a model lacking its trajectory learning, and it completed 10 out of 10 trials in a pig trachea.
That matters because endoscopy is one of the most common medical procedures in the world. Making the tool smarter about what the doctor intends could reduce procedure time, improve inspection quality, and lower the chance of complications from going the wrong way. It's still in the research stage, but it shows a future where medical instruments understand spoken words, not just a fixed button push.
- AI that follows spoken instructions could make endoscopes easier and safer to control.
- The system reduced wrong-direction movements by 83% compared to earlier designs.
- It worked in simulated organs and 10/10 tests on real animal tissue, inching closer to clinical use.
Why It Matters
Safer, more reliable internal exams with fewer mistakes—potentially lowering risk in routine colonoscopies and other procedures.