Do You Need to Move to See the World?
The Man Who Saw the World Upside Down

In 1896 the American psychologist George Stratton (1865–1957) strapped a set of lenses to his right eye that flipped the retinal image 180°. For three days everything looked upside down. When he reached for a cup, his hand shot off in the wrong direction. Walking felt sickening because the visual world swung wildly with every head turn. But after a few days something remarkable happened: the world began to feel right side up again. Stratton could pour tea, dodge a looming ball, and navigate his room. When he finally removed the lenses the real world briefly seemed alien—but it quickly snapped back to normal.
Stratton’s experiment raises a deep question: how does moving your body change what you see, and why did his vision adapt at all? Philosophers and scientists have argued for centuries about whether action—touch, eye movements, even the plan to move—is what gives our visual world its three‑dimensional shape. The answer touches everything from virtual reality to the way a newborn sees its mother’s face.
Berkeley’s Bold Claim: Vision Is Flat Until Touch Teaches It

George Berkeley (1685–1753), an Irish bishop and philosopher, kicked off the modern debate with a startling idea. In his New Theory of Vision (1709) he argued that the immediate objects of sight are just a two‑dimensional mosaic of light and color. Distance, solid shape, and outward depth are not seen directly. Instead, they are suggested by visual cues because those cues have been habitually linked to experiences of touch and movement.
Think of it like a language. The patch of blue and green that hits your eye when you look at a tree is like a word. You have learned that when you walk toward such a patch you will eventually feel rough bark under your fingers and the strain of your legs covering distance. Berkeley wrote that visible ideas are “the language whereby the governing spirit informs us what tangible ideas he is about to imprint upon us, in case we excite this or that motion in our own bodies.” Vision, on this view, is a forecast of tactile consequences. The spatial meaning of sight is borrowed from the body’s past actions.
Berkeley also pointed to eye‑muscle sensations as part of the story. When you focus on a nearby object, your eyes converge and the lens thickens; those muscular feelings become associated with close‑up tangible experiences. So depth perception depends on learning the associations between visual patterns and the patterns of movement and touch that go with them.
Critics were quick to object. Many said introspection shows no such train of tactile “images” popping up when we see distance—we just see a visible quality and nothing more. Others noted that sight often dominates touch. Look at a realistic painting of a curved vase on a flat canvas; your eyes insist the vase bulges outward even though your fingers know it is smooth. Condillac (1715–1780) captured this: “I have touched it, and yet this knowledge… does not prevent me from seeing convex figures.” And then there were baby animals—chicks that peck at grain minutes after hatching, without any time for a touch‑based education. If sight needs tutoring from touch, how do they manage?
Defenders of Berkeley could reply that the connection between sight and touch might be inborn in some creatures, not learned. But the cracks in the theory were widening. By the 19th century many philosophers and scientists began to suspect that movement enters vision even deeper—not just for depth, but for the basic layout of the two‑dimensional visual field itself.
How Your Eyes Move to Give You Direction

Hermann Lotze (1817–1881) and Hermann von Helmholtz (1821–1894) took the action‑based story a crucial step further. They asked: how does a point of light on your retina get tagged as “to the left” or “above”? The retina is just a sheet of cells; by itself it carries no direction labels. Lotze proposed that every spot on the retina has a local sign—a unique qualitative “feel” that tells the brain where that spot is. And what provides that feel? The answer, he said, is the pattern of kinaesthetic sensations you get when you move your eye to look directly at the spot.
Imagine a distant star that lands on a patch of your retina called P. To foveate it—to bring it into the center of gaze—your eye must travel a certain arc. Even if your eye does not actually move, stimulating P awakens a tiny muscular sensation that belongs to that arc, and that sensation recalls from memory the whole sequence of feelings you would have if you did move. Every location you see is thus tagged by the eye‑movement it would take to look at it. Helmholtz modified the idea: he argued that the relevant signal is not sensory feedback from the muscles, but the effort of will—the efference copy of the motor command sent out to move the eyes. In other words, the brain uses a copy of the order “move eye 10° right” to determine that the object currently stimulating the retina must be 10° left of where you are pointing your gaze.
This efference copy also solves a daily miracle: why the world does not jump every time you make a saccade. When you flick your eyes across a room, the retinal image sweeps drastically, yet the scene stays perfectly still. Helmholtz realized that the visual system compares the predicted displacement (from efference copy) with the actual shift on the retina. If they match, the perceived world is stable. If they mismatch—for example, when you push gently on your eyelid and the retinal image moves without a motor command—the world appears to lurch. This became known as the reafference principle: the brain subtracts out self‑caused sensory changes so that only externally caused changes show up in awareness.
Seeing Like a Skill: The Enactive Approach

In the early 2000s, J. Kevin O’Regan and Alva Noë pushed the action‑dependence of vision even further. Their enactive approach says that to perceive an object’s spatial properties is to have practical mastery of sensorimotor contingencies—the rules that govern how sensory input changes when you move your body. Seeing a cup’s shape is not just receiving an image; it is knowing that if you tilt your head to the left, the cup’s outline will transform in a certain law‑like way, while if you walk around it the handle will occlude part of the rim.
Their favorite piece of evidence comes from tactile‑visual sensory substitution devices (TVSS). A blind person wears a camera that converts video into a pattern of vibrations on the skin of the back. At first, the subject feels only tingles. But when the person is allowed to move the camera actively—panning, tilting, zooming—the sensations soon change character. Within hours many subjects report quasi‑visual experiences of objects arrayed out in front of them; they can duck a ball, judge distance, and recognize shapes. Crucially, if the camera is moved by someone else and the subject receives the same visual input passively, they learn nothing—the sensations stay meaningless buzzes. O’Regan and Noë argue that active movement is what allows the brain to extract the sensorimotor laws of the new “prosthetic” modality.
The enactive approach also claims that what you directly see is a two‑dimensional perspectival shape (P‑shape)—the flat outline an object projects onto a plane perpendicular to your line of sight—and that you understand the object’s true three‑dimensional shape only because you implicitly know how that P‑shape would change with movement. In that sense, the enactive view shares some DNA with Berkeley’s idea that vision without action is flat.
But challenges abound. Experiments on prism adaptation show that people can adjust to optical shifts even when they are moved passively in a wheelchair, without self‑produced movement—as long as they get clear information about the conflict between sight and the felt position of their limbs. The active movement may simply make the conflict more noticeable, not be necessary. And many vision scientists object that we do not first see a flat P‑shape and then infer solidity; instead, the visual system directly recovers three‑dimensional structure from cues like shadow, texture, and binocular disparity. A tilted coin looks like a tilted disk, not like a flat ellipse that we mentally reinterpret. Still, the enactive idea that perception is a skilled, action‑soaked activity has reshaped how we think about the mind’s connection to the world.
Why Your Brain Needs Action Plans

So where do these centuries of debate leave us? One influential synthesis comes from the disposition theory championed by Gareth Evans (1946–1980) and refined by Rick Grush. On this view, the spatial content of perception is neither purely visual nor purely motor. Instead, an organism’s perceptual and motor systems together form a single behavioral space. A sensory input counts as spatial—as placing an object there relative to your body—because it is woven into a rich network of dispositions to act. To see a mug as “straight ahead” is to have a set of fine‑grained readiness states: if you wanted to grasp it, your arm would move this way; if you wanted to look directly at it, your eyes would turn that amount. The “space” you experience is the common currency in which your brain maps both what you sense and what you can do.
The disposition theory does not require that you actually move—only that your brain is poised to issue the detailed commands that would produce the right actions in the right circumstances. This matters for many real‑life situations, including people who are paralyzed but still have vivid spatial awareness, and for the design of prosthetic devices and virtual reality. Your VR headset tricks you because it feeds your brain the exact sensorimotor contingencies that you have mastered in the real world: turn your head left, and the virtual scene flows right just as it should. When those rules break, you feel dizzy. When you adapt, like Stratton, your body has recalibrated the link between action and sensation.
The deepest lesson is that perception is not a passive movie screen. From Berkeley’s touch‑based language to the efference copies that steady your world, from sensorimotor rules to readiness to act, your own capacity for movement is part of what it means to see. You do not just observe space—you inhabit it by being a creature that acts.
Think about it
- If you put on glasses that flip your vision today, would you adapt more quickly than Stratton did? Why might someone who plays a lot of video games adapt differently?
- Imagine a robot that has perfect cameras but cannot move. Could that robot ever see space the way you do, or would it just have a flat picture?
- When you walk into a new room, you instantly know where the door is behind you without looking. Is that knowledge purely visual, or does it rely on your history of moving through similar rooms?





