Seeing What the Surgeon Cannot See
A surgeon can be looking directly at the surgical field and still not see the structure that matters most. Behind a thin wall of bone may sit an optic nerve, an internal carotid artery, the brain, or another critical structure separated from the surgeon by only a few millimeters. CT and MRI show those structures, but the endoscope can only show what lies within its visual field. That creates one of the fundamental challenges of minimally invasive skull-base surgery: the anatomy that is visible and the anatomy that the surgeon needs to understand are not always the same.
For decades, surgical navigation has helped bridge that gap. The endoscope provides the surgeon’s primary view while a separate navigation platform uses preoperative imaging and tracking technology to determine where the surgeon and instruments are positioned within the patient. The surgeon then has to continuously integrate those two representations of reality: what the camera is showing and what the imaging says lies beyond it.
Researchers at Brigham and Women’s Hospital and collaborating institutions are exploring whether that relationship could be redesigned. Their 2024 work, published in JAMA Otolaryngology–Head & Neck Surgery, tested whether a stereoscopic 3D endoscope could become more than a visualization device. By combining 3D endoscopic video with computer vision, simultaneous localization and mapping, or SLAM, and patient-specific CT and MRI models, the researchers demonstrated a potential pathway toward navigation in which the endoscope itself becomes a sensor for determining where its view exists within the patient’s anatomy.
The Camera Learns Where It Is
The headline might be 3D endoscopy, but the larger idea is spatial awareness. A conventional endoscope answers, “What am I looking at?” A navigation system answers, “Where am I?” The Brigham team’s approach begins to connect those questions by using stereoscopic vision to estimate depth and SLAM to track the endoscope as it moves through the surgical field.
The system combines overlapping views into a three-dimensional reconstruction of the exposed anatomy. Computer vision can then estimate the endoscope’s position, while registration aligns the reconstructed surface with a patient-specific model created from CT or MRI. Together, those steps connect what the surgeon sees with anatomical structures that may remain outside the immediate visual field.
The researchers evaluated the approach in nine deceased-donor specimens using an endoscopic endonasal approach to the anterior skull base. The exposed posterior sphenoid sinus was reconstructed from stereoscopic video and registered to CT-based three-dimensional models containing critical structures, including the internal carotid arteries and optic nerves. Mean reconstruction error was approximately 0.60 millimeters, mean registration error was 1.11 millimeters, and endoscope localization error was approximately 1.01 millimeters.
They also tested the concept retrospectively using surgical video from a patient who underwent endoscopic endonasal surgery for a tubercular meningioma. The reconstructed sphenoid sinus surface was registered to a model created from preoperative CT and MRI, producing a root-mean-square error of approximately 1.38 millimeters.
The results demonstrate technical feasibility, but they do not represent a clinically validated navigation system. The patient evaluation was retrospective, and the experiments did not establish that millimeter-level accuracy could be maintained throughout a complete operation as anatomy changes. What the study demonstrated was that endoscopic video could be transformed into a spatial reconstruction and connected to a larger patient-specific anatomical model. In other words, the study showed that the spatial relationship could be constructed, not that it could yet be trusted throughout a live operation.
When the Map Stops Being Trustworthy
That distinction becomes important because surgery is not a static environment. A cadaver provides a controlled setting, while a living operation introduces bleeding, moving instruments, tissue deformation, and progressive removal of bone and tissue. The surgeon is changing the environment at the same time the computer is trying to map it.
That creates a challenge for SLAM and other computer-vision approaches. The system relies on visual features that can be tracked as the endoscope moves, but those features can disappear when tissue is removed or become difficult to identify when blood or instruments obscure the field. At the same time, preoperative CT and MRI describe anatomy before surgery begins, while the endoscope is observing anatomy that may already have changed.
The researchers identified this limitation in their study. The cadaveric workflow used imaging obtained after surgical exposure, whereas clinical navigation normally relies on preoperative imaging. The work also did not reproduce many of the conditions encountered during live surgery, leaving open questions about how the system would perform as the surgical field evolves.
The challenge, then, may not simply be making the map more accurate. It may be knowing when the map can no longer be trusted. That changes the problem from simply asking, “Where am I?” to asking, “How certain am I that I still know where I am?” For surgical navigation, that distinction could be as important as the underlying localization accuracy.
From Navigation to Surgical Intelligence
The researchers also point toward a broader possibility: using the spatial model not only to help surgeons reach targets, but to identify structures they need to avoid. They describe the potential for machine learning to recognize instruments and critical anatomy and create safety boundaries around structures such as the optic nerves and internal carotid arteries. These boundaries are referred to as antitargets.
The concept is important because surgery is defined as much by avoidance as it is by targeting. A tumor may be the structure the surgeon needs to reach, while the carotid artery or optic nerve represents anatomy that must be protected. A future system could potentially combine patient-specific anatomical models with real-time endoscope localization and instrument tracking, identifying when an instrument approaches a predefined danger zone.
That capability remains prospective. The 2024 study did not demonstrate a clinical instrument-warning system. It established a technical foundation that could potentially support one.
The significance of this research, then, is not simply another advance in 3D endoscopy. It is the possibility of turning the endoscope from a device that shows anatomy into a system that understands its spatial relationship to that anatomy.
The endoscope began as a way to see inside the body through a small opening. Three-dimensional optics gave it depth. Computer vision can give it a sense of position. Medical imaging provides the larger anatomical context.
The next step could be a system that understands how those pieces relate and recognizes when that understanding is becoming unreliable.
The goal is not to replace the surgeon’s eyes, hands, or judgment. It is to give them a better representation of the environment in which they are operating, including critical anatomy they cannot directly see.
The endoscope may be learning where it is. The more important question is whether, someday, it will also know when it is wrong.
Learn More & Resources
To dive deeper into the technological convergence of 3D vision, simultaneous localization and mapping (SLAM), and advanced spatial intelligence in skull-base surgery, explore the primary research online:
Review the foundational open-access study on trackerless surgical navigation in the anterior skull base through PubMed Central.
References
Bartholomew, R. A., Zhou, H., Boreel, M., Suresh, K., Gupta, S., Mitchell, M. B., Hong, C., Lee, S. E., Smith, T. R., Guenette, J. P., Corrales, C. E., & Jagadeesan, J. Surgical navigation in the anterior skull base using 3-dimensional endoscopy and surface reconstruction. JAMA Otolaryngology–Head & Neck Surgery, 2024;150(4):318–326.






