The retina receives a flat, noisy, ambiguous image. What we experience is a stable three-dimensional world of objects. This note works through how that gap is closed: grouping, recognition, depth, constancy, and the decision rule that turns sensory evidence into a "yes, I saw it."
8 sections · 6 interactive labs · 7 PYQs · Cognitive Psychology
Sensation is how sensory receptors convert physical energy into neural signals: light hitting the retina, sound waves vibrating the cochlea, pressure activating skin receptors. Perception is the process of organising and interpreting that input to form a meaningful picture of the world. The same physical input can produce different percepts depending on context, expectation, and past experience.
Perception is constructive: the brain does not passively record the world but actively builds a representation of it. A flat retinal image gives rise to a rich three-dimensional percept. This gap between what the eye receives and what we actually experience is the central problem of perceptual psychology.
Neither account alone is enough. Modern models are interactive: top-down and bottom-up processes run simultaneously and constrain each other. Rumelhart's (1977) Interactive Activation model showed that letter-level and word-level information feed into each other, explaining the Word Superiority Effect.
The nervous system registers stimuli presented below the absolute threshold of conscious awareness. The standard method is priming: a subliminal stimulus is presented, and its influence on a later response is measured. Subliminal stimuli can shift attitudes but cannot produce complex learning.
A predisposition to perceive stimuli in a particular way, shaped by expectation, motivation, and past experience. Bruner and Minturn's (1955) classic demonstration: the same ambiguous figure is read as "B" in an alphabetic context and "13" in a numeric one.
Max Wertheimer, Wolfgang Köhler, and Kurt Koffka, working in Germany in the early 20th century, challenged the structuralist view that perception is built by adding up elementary sensations. Their central claim: "The whole is different from the sum of its parts." The key question was: what rules cause the visual system to organise elements into structured wholes?
Their answer was a set of principles of perceptual grouping: tendencies that cause elements to be grouped together based on specific properties. These principles operate automatically and preattentively. The overarching principle governing all of them is Prägnanz.
The visual system always organises stimuli into the simplest, most stable, and most regular form possible. Every other Gestalt law is a specific application of this principle.
Proximity. Elements close together in space tend to be grouped. Distance determines grouping before other properties: the same 36 dots now read as three column-pairs.
Before grouping laws operate, the visual system first divides any scene into a figure (the foreground object) and a ground (the background). This is not a grouping law but a prior assignment step that determines which region will be treated as an object. Rubin's Vase (1915) is the classic demonstration: the same contour supports two interpretations (two faces or a vase), and only one can be figure at a time.
The vase lies in front with a definite shape; the faces have receded into ground.
| Researcher | Year | Contribution | Key Term |
|---|---|---|---|
| Wertheimer | 1912 | Discovered apparent motion (phi phenomenon); this observation launched Gestalt psychology | Phi phenomenon; optimal ISI: 30–200 ms |
| Wertheimer | 1923 | Formal statement of grouping principles | Proximity, Similarity, Continuity, Closure, Common Fate |
| Rubin | 1915 | Figure-ground perception and reversible figures | Rubin's Vase |
| Köhler | Insight learning; critique of structuralism | Insight | |
| Koffka | 1935 | Applied Gestalt to development | Principles of Gestalt Psychology |
APPARENT MOTION · one dot seems to sweep across the gap
The Gestalt principles are descriptive rather than explanatory: they tell us what the visual system does, not how or why. The law of Prägnanz is also circular ("simple" is partly defined by what we already perceive as simple). More recent work grounds these principles in the statistical regularities of natural scenes.
How do we recognise that a pattern represents a particular object or letter? A handwritten "A" and a printed "A" look physically different, yet we recognise both instantly. Several competing theories have emerged, organised here by processing level.
The visual system stores a complete template (a mental copy) for every known pattern and recognises an incoming stimulus by finding the closest match. Template theory fails in practice. It requires storing a separate template for every size, orientation, and style of every object. It cannot explain recognition of novel instances, and it predicts that rotation should strongly impair recognition, which it does not.
Objects are not stored as wholes but as lists of features: lines, edges, angles, and curves. Recognition occurs when the incoming features match a stored feature list. Hubel and Wiesel (1959–1962) found that neurons in primary visual cortex respond selectively to specific features: simple cells to oriented lines, complex cells to oriented edges anywhere in a region, and hypercomplex cells to corners and moving edges.
Oliver Selfridge's (1959) Pandemonium Model formalised this computationally: image demons hold raw input, feature demons each respond to one feature, cognitive demons shout when their features are detected, and the decision demon picks the loudest cognitive demon. This hierarchical model anticipated convolutional neural networks.
Features handle size and position variation naturally: the same features appear wherever on the retina the object falls and however large it is. Feature theory also explains confusion errors: E and F (many shared features) are far more often confused than X and O (few shared features). Templates cannot account for this pattern.
The visual system decomposes the image through a bank of spatial frequency channels: neural filters each tuned to a different range of coarse-to-fine detail, like a Fourier analysis. Low frequencies carry global shape; high frequencies carry fine edges. After prolonged viewing of one spatial frequency, sensitivity to that frequency drops selectively (channel fatigue) while other frequencies are unaffected. This selective fatigue is strong evidence for independent, parallel channels.
David Marr proposed three successive representations for the visual system:
Object recognition ultimately requires a viewpoint-independent representation: a chair must be recognisable from any angle. The 3-D model achieves this.
Irving Biederman proposed a specific vocabulary for the 3-D model: geons, a set of about 36 simple volumetric primitives (cylinders, cones, blocks, wedges). Objects are represented as structured geon assemblies. The visual system parses the retinal image into its component geons and matches this description against stored assemblies.
Geons are viewpoint-invariant: they are identifiable from almost any angle. Degrading the junctions where geons meet impairs recognition far more than degrading the middle portions, because junctions carry the structural description needed for part-based parsing.
RBC relies on viewpoint-independent geon descriptions. View-based models (Tarr and Bülthoff; Poggio) propose that we store multiple 2D snapshots and recognise objects by interpolating between them. The debate: is recognition truly viewpoint-independent, or does rotating to an unfamiliar angle cost something? Evidence supports view-based accounts for many objects, especially novel ones; geon-based processes may dominate for familiar objects with clear part structure.
Gibson rejected the computational framework. Form recognition does not involve matching features or assembling geons. The visual world provides invariant information in the optic array that the visual system picks up directly, without computation, inference, or stored templates. Objects are perceived as wholes, specified by their relationship to the surrounding optic array. This is a strongly bottom-up, anti-representational position.
Probabilistic Functionalism describes how observers use multiple imperfect cues to infer true properties of distal objects. Each cue has an ecological validity (how reliably it correlates with the real property in nature) and a utilisation weight (how much the perceiver relies on it). Accurate perception, in this model, reflects calibration of cue weights to their ecological validity over time. This is a probabilistic counterpart to Gibson's direct pickup view: both are ecology-grounded, but Brunswick sees cue-weighting where Gibson sees direct specification.
| Theory | Main Idea | Strength | Limitation | Theorist |
|---|---|---|---|---|
| Template Matching | Store complete copies; compare stimulus to all stored templates | Simple and intuitive | Needs an enormous template store; fails with novel instances | |
| Feature Detection | Objects described by feature lists; recognise by feature matching | Neurophysiologically grounded; handles size and position variation | Ignores spatial relations between features | Hubel & Wiesel; Selfridge |
| Spatial Frequency Channels | Visual system uses a bank of spatial frequency filters | Strong psychophysical and neural support | Hard to link directly to high-level object recognition | Campbell & Robson (1968) |
| Computational (Marr) | Three-stage representation: Primal Sketch, 2½-D Sketch, 3-D Model | Formally specifies the problem at multiple levels | 3-D model construction underspecified; view-based evidence challenges it | Marr (1982) |
| Recognition by Components | Objects are geon assemblies; recognise by matching structural descriptions | Viewpoint invariance; explains part-based recognition | Struggles with within-category discrimination (all faces have the same geons) | Biederman (1987) |
| Direct Perception | Forms directly specified by optic array invariants; no templates or computation needed | Explains accurate real-world perception without cognitive overhead | Does not explain recognition of degraded or impoverished stimuli | Gibson (1979) |
Apperceptive agnosia: the patient cannot copy, match, or discriminate shapes even though basic sensory processing is intact (perceptual stage fails). Associative agnosia: perceptual processing is intact (the patient can draw an object accurately) but they cannot name or identify what they drew (semantic disconnection). Prosopagnosia: face-specific agnosia, typically from damage to the fusiform face area.
The retina is a two-dimensional surface, yet we perceive a three-dimensional world. The visual system recovers depth using two classes of cues: monocular cues (available to one eye) and binocular cues (requiring both eyes).
Monocular cues are also called pictorial cues because they convey depth in a flat picture. Artists have exploited them for centuries.
Linear perspective. Parallel lines appear to converge toward a vanishing point as they recede into the distance: the road edges meet at the horizon.
| Cue | How it works | Example |
|---|---|---|
| Linear Perspective | Parallel lines appear to converge toward a vanishing point as they recede into the distance | Railway tracks appearing to meet at the horizon |
| Texture Gradient | Surface texture becomes finer and more densely packed with increasing distance | A cobblestone path with stones appearing smaller in the distance |
| Interposition (Occlusion) | When one object blocks part of another, the occluding object is perceived as closer | A house partially hidden behind a tree |
| Relative Size | When two similar-sized objects are compared, the smaller retinal image signals greater distance | Two identical cars at different distances |
| Height in Visual Field | Objects higher in the visual field (for ground-level surfaces) are perceived as more distant | Mountains appearing above foreground objects near the horizon |
| Aerial Perspective | Distant objects appear hazy, less saturated, and bluish due to atmospheric light scattering | Distant mountains appearing blue-grey |
| Motion Parallax | As the observer moves, nearby objects sweep across the visual field faster than distant ones | Nearby trees fly past a train window while distant hills barely move |
| Accommodation | The ciliary muscle adjusts lens curvature for near vs. far objects; the muscle tension provides a distance signal (effective within about 2 m) | Reading at arm's length vs. looking across a room |
The eyes are about 6.5 cm apart, so each receives a slightly different image. The visual system uses this difference to compute depth, a process called stereopsis. Disparity is greatest for nearby objects and decreases with distance.
When fixating a near object, both eyes rotate inward. The degree of inward rotation gives the brain a muscular distance signal. Effective for objects within about 6 m.
Gibson argued that depth cues are not inferred by the visual system. They are directly specified by the structure of the optic array. He emphasised texture gradients as invariant properties of ambient light that directly specify surface layout, requiring no computation or prior knowledge.
Perceptual constancy is the ability to perceive objects as having stable properties (size, shape, colour, brightness) despite continuous changes in the retinal image caused by changes in distance, angle, and illumination. Without constancy, a friend walking away would appear to shrink, and a tilted plate would appear elliptical.
We perceive objects as having stable physical size even as their retinal image changes with distance. The visual system rescales retinal size by factoring in perceived distance. This is the Size-Distance Invariance Hypothesis: Perceived Size = Retinal Image Size × Perceived Distance. If perceived distance is miscalculated, perceived size is distorted. This distortion is the basis of many visual illusions.
After fixating a coloured square, look at a near wall: the afterimage appears small. Look at a far wall: it appears large. Retinal size is constant; what changes is perceived distance. Perceived size scales directly with perceived distance, confirming size-distance invariance.
| Type | What stays constant | Despite changes in |
|---|---|---|
| Shape constancy | Perceived shape of an object | Viewing angle (a door swinging open still looks rectangular) |
| Colour constancy | Perceived colour of surfaces | Illumination (a red apple looks red under sunlight and artificial light) |
| Brightness constancy | Perceived lightness of surfaces | Overall illumination level (a white page looks white even in dim light) |
| Location constancy | Perceived position of stationary objects | Eye and head movements that shift the retinal image |
Perceptual adaptation is the ability of the perceptual system to recalibrate when it receives systematically distorted input. Stratton (1897) had participants wear goggles that inverted the visual field. Initially this caused disorientation. Over several days, participants adapted and their perception normalised. The key finding: active movement (physically interacting with the environment) drove adaptation, not passive exposure alone. The perceptual system recalibrates through sensorimotor feedback.
Visual illusions are cases where the percept systematically diverges from physical properties. They are valuable because they reveal the assumptions the visual system uses. An illusion is a case where those assumptions fail.
| Illusion | Phenomenon | Explanation |
|---|---|---|
| Müller-Lyer | Two equal lines look different: outward arrowheads make a line look longer | Arrowheads create depth cues resembling near and far room corners. Size constancy scaling is applied to a flat figure, distorting length. |
| Ponzo | Two identical bars between converging lines look different in size | Converging lines provide a linear perspective depth cue. The upper bar appears farther, so constancy scaling makes it look larger. |
| Ebbinghaus (Titchener) | A circle surrounded by small circles appears larger than the same circle surrounded by large ones | Contrast effect: the central circle is judged relative to surrounding context, not by absolute retinal size. |
| Moon Illusion | The moon looks larger near the horizon than high in the sky, despite identical angular size | The horizon moon appears farther (terrain depth cues intervene), so constancy scaling makes it look larger. The zenith moon lacks these cues. |
| Ames Room | People at opposite corners of a distorted room look radically different in size | The room fools the visual system into assuming a rectangular shape. When that assumption is wrong, constancy scaling produces large distortions. |
| Necker Cube / Rubin's Vase | Ambiguous figures that spontaneously reverse between two interpretations | Both interpretations fit the input equally well. The visual system generates competing hypotheses and alternates between them (Gregory). |
Many geometric illusions arise because the visual system misapplies size-constancy scaling to flat, two-dimensional figures. Cues in the figure (converging lines, arrowheads) trigger depth-processing mechanisms appropriate for 3D scenes; applied to a flat figure, they produce distortions. Gregory predicted that people with less exposure to carpentered (right-angle-rich) environments should show weaker illusions, a prediction with partial support (Segall, Campbell and Herskovits, 1966).
A stationary object appears to move when its surrounding frame moves. The visual system uses the frame as a reference point and assigns relative motion to the smaller enclosed object. Classic example: when the train beside you moves, your stationary train appears to move in the opposite direction.
Adaptation level (AL) is a weighted average of all stimuli of a given class that the observer has recently been exposed to. Stimuli above the AL appear intense or large; those below appear weak or small; those at the AL appear neutral. The AL shifts with exposure: after lifting heavy weights, moderate weights feel lighter. This context-dependence is the same mechanism behind contrast-based illusions, where a stimulus looks different depending on what surrounds it.
In normal conversation, speakers produce about 10–15 phonemes per second. Listeners parse this stream effortlessly, recognise words, distinguish speakers, and track meaning simultaneously. Several features of the speech signal make this difficult to explain.
There is no acoustic invariance: the same phoneme sounds different depending on surrounding sounds (coarticulation), the speaker's voice, rate, and accent. There are no clear word boundaries in the continuous stream, yet listeners hear discrete units. Context also provides powerful top-down constraints: when one phoneme is replaced by a cough, listeners hear the complete word without noticing the gap. Warren (1970) called this the Phoneme Restoration Effect.
| Theory | Main Claim | Mechanism | Strength | Limitation |
|---|---|---|---|---|
| Motor Theory (Liberman et al., 1967) | Listeners recover the speaker's intended articulatory gestures, not acoustic patterns | A speech-specific module uses knowledge of one's own articulation to decode what the speaker intended | Explains categorical perception and the McGurk effect | Infants and people with severe motor disorders still perceive speech normally |
| TRACE Model (McClelland & Elman, 1986) | Three interactive levels: features, phonemes, and words. Activation flows both upward and downward | Word-level nodes feed back to activate expected phonemes, which activate expected features. This two-way flow explains context effects naturally | Simulates human data well; explains phoneme restoration and the Ganong Effect | Computationally complex; timing of top-down feedback is debated |
| Cohort Model (Marslen-Wilson & Tyler, 1980) | Initial phonemes activate a "cohort" of all matching words, which narrows as more input arrives | Context and word frequency prune the cohort further until one word remains | Explains why listeners often identify a word before it finishes (word-onset effect) | If the first phoneme is misheard, the target word never enters the cohort (addressed in the 1987 revision) |
| Direct Realist Theory (Fowler, 1986) | Listeners directly perceive the articulatory gestures that produced the sound, without a special module | Articulatory gestures are real events in the world; the acoustic signal is structured by them, so listeners evolved to perceive these events directly | Avoids a special speech module; integrates with Gibson's general ecological approach | Vague about how listeners "directly" perceive gestures from sound waves |
Traditional psychophysics assumed a fixed absolute threshold: below it, stimuli are never detected; above it, they always are. Signal Detection Theory (SDT) replaced this. Detection depends on two independent factors: the observer's sensitivity (d-prime) and their response criterion (beta), which is how cautious or liberal they are about reporting a signal. The same sensitivity can produce very different detection rates depending on the criterion.
Measures how far apart the signal-plus-noise and noise-alone distributions are. A higher d-prime means the observer can more reliably tell signal from noise. Sensitivity is a true perceptual measure, independent of motivation or strategy.
The decision cut-off: how much evidence the observer requires before saying "yes." When missing a signal carries severe consequences, the observer lowers the criterion, accepting more false alarms to avoid misses. Criterion is motivational, not perceptual.
| Signal Present | Signal Absent | |
|---|---|---|
| Respond "Yes" | Hit | False Alarm |
| Respond "No" | Miss | Correct Rejection |
A radiologist reading chest X-rays knows that missing a tumour is catastrophic, so they lower the criterion and report any hint of a shadow. A security guard faces the opposite trade-off (false alarms waste searches) and raises the criterion. Both may have identical d-prime but different hit rates because their criteria differ. SDT separates this motivational factor from true perceptual sensitivity.
Phi-Phenomenon is best seen between which of the following time intervals?
Phi-Phenomenon is best seen between which of the following time intervals?
Among the laws of perceptual grouping, the law of simplicity is a tendency to:
A man judged to be six feet tall when standing at ten feet away has a retinal image of size X. At twenty feet, the retinal image is X/2. How tall shall he be perceived at a distance of five feet?
Match the following: List I (Name): a. Template Matching Model b. Feature Matching Model c. Recognition by Components Model d. Configuration Model; List II (Feature): i. Spatial relations deviate from prototype ii. 3D objects described via parts and spatial relations iii. Visual analysis detects colours and edges iv. Recognition of barcodes.
Signal detection depends upon:
An important factor which enables one to adapt to inverted vision is:
Some people are able to draw an object, match similar objects and describe its component parts, but fail to recognize what they have just seen or drawn. This describes:
| Theorist(s) | Year | Contribution | Key Term |
|---|---|---|---|
| Wertheimer | 1912, 1923 | Founded Gestalt psychology; phi phenomenon; grouping laws | Apparent Motion, Gestalt Principles |
| Rubin | 1915 | Figure-ground perception and reversible figures | Rubin's Vase, Figure-Ground |
| Gibson, J.J. | 1950, 1979 | Ecological approach; optic array; affordances; texture gradients | Direct Perception, Affordances |
| Gregory | 1963, 1966 | Constructivist theory; misapplied constancy explanation of illusions | Perception as Hypothesis, Misapplied Constancy |
| Hubel and Wiesel | 1959–1962 | Feature-detecting neurons in primary visual cortex | Simple, Complex, Hypercomplex Cells |
| Selfridge | 1959 | Pandemonium model of pattern recognition | Feature Demons, Cognitive Demons |
| Campbell and Robson | 1968 | Multiple spatial frequency channels in vision | Spatial Frequency, Channel Adaptation |
| Marr | 1982 | Computational theory of vision; three-stage representation | Primal Sketch, 2½-D Sketch, 3-D Model |
| Biederman | 1987 | Recognition by Components (RBC) | Geons, Viewpoint Invariance |
| Emmert | 1881 | Size-distance invariance; afterimage size scaling | Emmert's Law |
| Brunswick | 1956 | Probabilistic functionalism; lens model of perception | Ecological Validity, Cue Utilisation |
| Helson | 1964 | Adaptation-level theory | Adaptation Level (AL) |
| Liberman et al. | 1967 | Motor theory of speech perception | Articulatory Gestures, Categorical Perception |
| McClelland and Elman | 1986 | TRACE model of speech perception | Interactive Activation, Top-Down Feedback |
| Marslen-Wilson and Tyler | 1980 | Cohort model of spoken word recognition | Cohort, Recognition Point |
| Fowler | 1986 | Direct Realist Theory of speech perception | Articulatory Events as Distal Objects |
| McGurk and MacDonald | 1976 | McGurk Effect: multimodal speech perception | Visual-Auditory Integration |
| Warren | 1970 | Phoneme Restoration Effect | Top-Down Lexical Completion |
| Segall, Campbell and Herskovits | 1966 | Cultural variation in susceptibility to geometric illusions | Carpentered World Hypothesis |
| Stratton | 1897 | Perceptual adaptation to inverted vision | Active Movement, Perceptual Recalibration |
| Bruner and Minturn | 1955 | Perceptual set ("B vs 13" study) | Perceptual Set, Context Effects |