Computer Vision
How machines build usable understanding from pixels — robust perception, visual reasoning, and the gap between recognizing and comprehending.
Read the deep dive →Independent AI Research Laboratory
AetherNeural is an independent research laboratory studying computer vision and agent enhancement — the meeting point where seeing becomes doing. We pursue long questions, publish open notes, and treat every failed experiment as data.
Research areas
Our work is organized around durable questions rather than product cycles. Each area feeds the others — perception informs action, and action reveals what perception was missing.
How machines build usable understanding from pixels — robust perception, visual reasoning, and the gap between recognizing and comprehending.
Read the deep dive →Making autonomous agents more capable and more trustworthy — planning, tool use, memory, and the craft of reliable long-horizon behavior.
Read the deep dive →Where sight meets language, sound, and action — how agents fuse modalities into a coherent picture of the world instead of parallel guesses.
Read the deep dive →Measuring what we claim to measure — rigorous, reproducible evaluation design for vision models and agents, including what benchmarks miss.
Read the deep dive →Featured projects
A selection of the experiments currently running in the lab. Each is documented with its open questions — including the ones we haven't answered.
Teaching agents to verify what they see before they act — grounding interface elements, documents, and scenes into checkable references instead of guesses.
Exploring how vision systems can express honest uncertainty about what they observe — and how downstream agents should respond when confidence is low.
What should an agent remember across a long task — and what should it deliberately forget? Studying episodic and semantic memory for sustained work.
Lab journal
Short-form observations from the bench — things we noticed, questions that surfaced, and mistakes worth writing down. Editorial, not peer-reviewed; honest, not inflated.
Giving a vision model a sharper image doesn't always give it better understanding. On why pixel count and comprehension diverge — and what actually helps.
Read the note →Failure patterns we've observed in agents working from screenshots — misread states, phantom buttons, and the quiet discipline of verifying before clicking.
Read the note →Most of our experiments don't work. That is the work. A short argument for writing down what failed, and reading other people's failures too.
Read the note →About the lab
AetherNeural exists because some questions are best pursued slowly, without a product roadmap attached. We are a small independent lab with no institutional affiliation — which means our only accountability is to the quality of the questions and the honesty of the answers.
Our story and principles →Collaborate
We welcome correspondence from researchers, builders, and the simply curious — on vision, agents, evaluation, or anything at their intersection.
Get in touch