Attention & Perception

Perceptual Organization

Cognitive PsychologyHigh

Study map

The retina receives a flat, noisy, ambiguous image. What we experience is a stable three-dimensional world of objects. This note works through how that gap is closed: grouping, recognition, depth, constancy, and the decision rule that turns sensory evidence into a "yes, I saw it."

8 sections · 6 interactive labs · 7 PYQs · Cognitive Psychology

Section 01

Sensation and Perception

IMP

Sensation is how sensory receptors convert physical energy into neural signals: light hitting the retina, sound waves vibrating the cochlea, pressure activating skin receptors. Perception is the process of organising and interpreting that input to form a meaningful picture of the world. The same physical input can produce different percepts depending on context, expectation, and past experience.

Perception is constructive: the brain does not passively record the world but actively builds a representation of it. A flat retinal image gives rise to a rich three-dimensional percept. This gap between what the eye receives and what we actually experience is the central problem of perceptual psychology.

Bottom-Up Processing

  • Starts with raw sensory data and builds upward to meaning
  • The stimulus drives the process; no prior knowledge is needed
  • Gibson's Direct Perception theory
  • Affordances specify what actions the environment makes possible
  • Works best in rich, natural environments where sensory data is full

Top-Down Processing

  • Knowledge, schemas, and expectations shape how raw input is interpreted
  • Prior experience, language, culture, and goals all contribute
  • Gregory's Constructivist Theory: perception as a hypothesis
  • Explains the word-superiority effect and priming
  • Works best under degraded or ambiguous conditions

Neither account alone is enough. Modern models are interactive: top-down and bottom-up processes run simultaneously and constrain each other. Rumelhart's (1977) Interactive Activation model showed that letter-level and word-level information feed into each other, explaining the Word Superiority Effect.

Subliminal Perception

The nervous system registers stimuli presented below the absolute threshold of conscious awareness. The standard method is priming: a subliminal stimulus is presented, and its influence on a later response is measured. Subliminal stimuli can shift attitudes but cannot produce complex learning.

Perceptual Set

A predisposition to perceive stimuli in a particular way, shaped by expectation, motivation, and past experience. Bruner and Minturn's (1955) classic demonstration: the same ambiguous figure is read as "B" in an alphabetic context and "13" in a numeric one.

Section 02

Gestalt Laws of Perceptual Organization

V. IMP

Max Wertheimer, Wolfgang Köhler, and Kurt Koffka, working in Germany in the early 20th century, challenged the structuralist view that perception is built by adding up elementary sensations. Their central claim: "The whole is different from the sum of its parts." The key question was: what rules cause the visual system to organise elements into structured wholes?

Their answer was a set of principles of perceptual grouping: tendencies that cause elements to be grouped together based on specific properties. These principles operate automatically and preattentively. The overarching principle governing all of them is Prägnanz.

Law of Prägnanz (Good Form)

The visual system always organises stimuli into the simplest, most stable, and most regular form possible. Every other Gestalt law is a specific application of this principle.

Lab 01 · Gestalt playgroundsame 36 dots, six laws

Proximity. Elements close together in space tend to be grouped. Distance determines grouping before other properties: the same 36 dots now read as three column-pairs.

Every principle above is a specific application of Prägnanz: switch tabs and watch the identical dots reorganise into the simplest, most stable grouping available.
Fig.Grouping principles, live — switch the rule and watch the same dots regroup
Proximity, Similarity, Continuity, Closure, Common Fate, and Symmetry acting on one identical dot field. Nothing changes except the organising rule.

Figure and Ground

Before grouping laws operate, the visual system first divides any scene into a figure (the foreground object) and a ground (the background). This is not a grouping law but a prior assignment step that determines which region will be treated as an object. Rubin's Vase (1915) is the classic demonstration: the same contour supports two interpretations (two faces or a vase), and only one can be figure at a time.

Lab 02 · Figure-ground flipRubin's Vase, 1915
vase is figure

The vase lies in front with a definite shape; the faces have receded into ground.

One contour, two objects, never both at once: the boundary always belongs to the figure, and whichever region loses it turns into shapeless ground.
Fig.Rubin's Vase — force the reversal yourself
The contour never moves. Only its ownership changes, and with it which region has shape.

Figure

  • Has a definite shape; appears to lie in front
  • Smaller, more symmetric, or more convex region
  • Surrounded, vertical, or familiar regions favour figure status
  • The contour between figure and ground belongs to the figure

Ground

  • Appears shapeless; extends behind the figure
  • Larger, less regular, less symmetric region
  • No clear boundary of its own
  • May appear to extend beyond the visual field

Key Gestalt Researchers

ResearcherYearContributionKey Term
Wertheimer1912Discovered apparent motion (phi phenomenon); this observation launched Gestalt psychologyPhi phenomenon; optimal ISI: 30–200 ms
Wertheimer1923Formal statement of grouping principlesProximity, Similarity, Continuity, Closure, Common Fate
Rubin1915Figure-ground perception and reversible figuresRubin's Vase
KöhlerInsight learning; critique of structuralismInsight
Koffka1935Applied Gestalt to developmentPrinciples of Gestalt Psychology
Lab 03 · Phi phenomenonWertheimer, 1912
ISI = 100 ms
5 ms30–200 ms · phi window350 ms

APPARENT MOTION · one dot seems to sweep across the gap

The exam anchor: optimal apparent motion at an interstimulus interval of 30–200 ms. Below 30 ms the flashes fuse; above 200 ms they fall apart into succession.
Fig.Phi phenomenon — drag the ISI slider through all three regimes
Below ~30 ms the two dots fuse into one; between 30–200 ms a single dot appears to move; above 200 ms you see two separate flashes. This is the exact interval asked in the UGC NET item below.

Critique

The Gestalt principles are descriptive rather than explanatory: they tell us what the visual system does, not how or why. The law of Prägnanz is also circular ("simple" is partly defined by what we already perceive as simple). More recent work grounds these principles in the statistical regularities of natural scenes.

Section 03

Form Recognition Theories

V. IMP

How do we recognise that a pattern represents a particular object or letter? A handwritten "A" and a printed "A" look physically different, yet we recognise both instantly. Several competing theories have emerged, organised here by processing level.

Early Theories

Template Matching

The visual system stores a complete template (a mental copy) for every known pattern and recognises an incoming stimulus by finding the closest match. Template theory fails in practice. It requires storing a separate template for every size, orientation, and style of every object. It cannot explain recognition of novel instances, and it predicts that rotation should strongly impair recognition, which it does not.

Feature Detection Theory

Objects are not stored as wholes but as lists of features: lines, edges, angles, and curves. Recognition occurs when the incoming features match a stored feature list. Hubel and Wiesel (1959–1962) found that neurons in primary visual cortex respond selectively to specific features: simple cells to oriented lines, complex cells to oriented edges anywhere in a region, and hypercomplex cells to corners and moving edges.

Oliver Selfridge's (1959) Pandemonium Model formalised this computationally: image demons hold raw input, feature demons each respond to one feature, cognitive demons shout when their features are detected, and the decision demon picks the loudest cognitive demon. This hierarchical model anticipated convolutional neural networks.

Why feature theory beats templates

Features handle size and position variation naturally: the same features appear wherever on the retina the object falls and however large it is. Feature theory also explains confusion errors: E and F (many shared features) are far more often confused than X and O (few shared features). Templates cannot account for this pattern.

Spatial Frequency Channels (Campbell & Robson, 1968)

The visual system decomposes the image through a bank of spatial frequency channels: neural filters each tuned to a different range of coarse-to-fine detail, like a Fourier analysis. Low frequencies carry global shape; high frequencies carry fine edges. After prolonged viewing of one spatial frequency, sensitivity to that frequency drops selectively (channel fatigue) while other frequencies are unaffected. This selective fatigue is strong evidence for independent, parallel channels.

Structural Theories

Marr's Computational Theory (1982)

David Marr proposed three successive representations for the visual system:

Primal SketchEdges, blobs, bars
2½-D SketchDepth, orientation, viewer-centred
3-D ModelObject-centred; viewpoint-independent

Object recognition ultimately requires a viewpoint-independent representation: a chair must be recognisable from any angle. The 3-D model achieves this.

Recognition by Components (Biederman, 1987)

Irving Biederman proposed a specific vocabulary for the 3-D model: geons, a set of about 36 simple volumetric primitives (cylinders, cones, blocks, wedges). Objects are represented as structured geon assemblies. The visual system parses the retinal image into its component geons and matches this description against stored assemblies.

CYLINDERCONEBLOCKWEDGE
Fig.The four basic geons
Geons are recovered from non-accidental properties of edges, which is why they survive changes in viewpoint.

Geons are viewpoint-invariant: they are identifiable from almost any angle. Degrading the junctions where geons meet impairs recognition far more than degrading the middle portions, because junctions carry the structural description needed for part-based parsing.

RBC vs. View-Based Models

RBC relies on viewpoint-independent geon descriptions. View-based models (Tarr and Bülthoff; Poggio) propose that we store multiple 2D snapshots and recognise objects by interpolating between them. The debate: is recognition truly viewpoint-independent, or does rotating to an unfamiliar angle cost something? Evidence supports view-based accounts for many objects, especially novel ones; geon-based processes may dominate for familiar objects with clear part structure.

Ecological Approach

Gibson's Direct Perception

Gibson rejected the computational framework. Form recognition does not involve matching features or assembling geons. The visual world provides invariant information in the optic array that the visual system picks up directly, without computation, inference, or stored templates. Objects are perceived as wholes, specified by their relationship to the surrounding optic array. This is a strongly bottom-up, anti-representational position.

Brunswick's Lens Model (1956)

Probabilistic Functionalism describes how observers use multiple imperfect cues to infer true properties of distal objects. Each cue has an ecological validity (how reliably it correlates with the real property in nature) and a utilisation weight (how much the perceiver relies on it). Accurate perception, in this model, reflects calibration of cue weights to their ecological validity over time. This is a probabilistic counterpart to Gibson's direct pickup view: both are ecology-grounded, but Brunswick sees cue-weighting where Gibson sees direct specification.

TheoryMain IdeaStrengthLimitationTheorist
Template MatchingStore complete copies; compare stimulus to all stored templatesSimple and intuitiveNeeds an enormous template store; fails with novel instances
Feature DetectionObjects described by feature lists; recognise by feature matchingNeurophysiologically grounded; handles size and position variationIgnores spatial relations between featuresHubel & Wiesel; Selfridge
Spatial Frequency ChannelsVisual system uses a bank of spatial frequency filtersStrong psychophysical and neural supportHard to link directly to high-level object recognitionCampbell & Robson (1968)
Computational (Marr)Three-stage representation: Primal Sketch, 2½-D Sketch, 3-D ModelFormally specifies the problem at multiple levels3-D model construction underspecified; view-based evidence challenges itMarr (1982)
Recognition by ComponentsObjects are geon assemblies; recognise by matching structural descriptionsViewpoint invariance; explains part-based recognitionStruggles with within-category discrimination (all faces have the same geons)Biederman (1987)
Direct PerceptionForms directly specified by optic array invariants; no templates or computation neededExplains accurate real-world perception without cognitive overheadDoes not explain recognition of degraded or impoverished stimuliGibson (1979)

Recognition failures: agnosias

Apperceptive agnosia: the patient cannot copy, match, or discriminate shapes even though basic sensory processing is intact (perceptual stage fails). Associative agnosia: perceptual processing is intact (the patient can draw an object accurately) but they cannot name or identify what they drew (semantic disconnection). Prosopagnosia: face-specific agnosia, typically from damage to the fusiform face area.

Section 04

Depth Perception

V. IMP

The retina is a two-dimensional surface, yet we perceive a three-dimensional world. The visual system recovers depth using two classes of cues: monocular cues (available to one eye) and binocular cues (requiring both eyes).

Monocular Cues to Depth

Monocular cues are also called pictorial cues because they convey depth in a flat picture. Artists have exploited them for centuries.

Lab 04 · One scene, six depth cuestap a cue to expose it

Linear perspective. Parallel lines appear to converge toward a vanishing point as they recede into the distance: the road edges meet at the horizon.

Every cue below is monocular: close one eye and the scene loses none of its depth. Artists have been packing all six into flat canvases for centuries.
Fig.One scene, six monocular cues — tap a chip to isolate its region
Every cue in the table below is present simultaneously in this single flat image. That is exactly why a picture can carry depth.
CueHow it worksExample
Linear PerspectiveParallel lines appear to converge toward a vanishing point as they recede into the distanceRailway tracks appearing to meet at the horizon
Texture GradientSurface texture becomes finer and more densely packed with increasing distanceA cobblestone path with stones appearing smaller in the distance
Interposition (Occlusion)When one object blocks part of another, the occluding object is perceived as closerA house partially hidden behind a tree
Relative SizeWhen two similar-sized objects are compared, the smaller retinal image signals greater distanceTwo identical cars at different distances
Height in Visual FieldObjects higher in the visual field (for ground-level surfaces) are perceived as more distantMountains appearing above foreground objects near the horizon
Aerial PerspectiveDistant objects appear hazy, less saturated, and bluish due to atmospheric light scatteringDistant mountains appearing blue-grey
Motion ParallaxAs the observer moves, nearby objects sweep across the visual field faster than distant onesNearby trees fly past a train window while distant hills barely move
AccommodationThe ciliary muscle adjusts lens curvature for near vs. far objects; the muscle tension provides a distance signal (effective within about 2 m)Reading at arm's length vs. looking across a room

Binocular Cues to Depth

Binocular Disparity

The eyes are about 6.5 cm apart, so each receives a slightly different image. The visual system uses this difference to compute depth, a process called stereopsis. Disparity is greatest for nearby objects and decreases with distance.

Convergence

When fixating a near object, both eyes rotate inward. The degree of inward rotation gives the brain a muscular distance signal. Effective for objects within about 6 m.

Gibson's Ecological Approach to Depth

Gibson argued that depth cues are not inferred by the visual system. They are directly specified by the structure of the optic array. He emphasised texture gradients as invariant properties of ambient light that directly specify surface layout, requiring no computation or prior knowledge.

Nativist vs. Empiricist on Depth Perception

Nativist

  • Depth perception is innate
  • The visual system is pre-wired to interpret retinal disparity as depth
  • Supported by research showing depth sensitivity in newborns
  • Associated with Descartes

Empiricist

  • Depth perception is learned through experience
  • We learn to associate visual cues with distance through tactile experience
  • Berkeley: no direct visual access to depth; it is always inferred
  • Associated with Berkeley
Section 05

Perceptual Constancy

IMP

Perceptual constancy is the ability to perceive objects as having stable properties (size, shape, colour, brightness) despite continuous changes in the retinal image caused by changes in distance, angle, and illumination. Without constancy, a friend walking away would appear to shrink, and a tilted plate would appear elliptical.

Size Constancy

We perceive objects as having stable physical size even as their retinal image changes with distance. The visual system rescales retinal size by factoring in perceived distance. This is the Size-Distance Invariance Hypothesis: Perceived Size = Retinal Image Size × Perceived Distance. If perceived distance is miscalculated, perceived size is distorted. This distortion is the basis of many visual illusions.

Emmert's Law (1881)

After fixating a coloured square, look at a near wall: the afterimage appears small. Look at a far wall: it appears large. Retinal size is constant; what changes is perceived distance. Perceived size scales directly with perceived distance, confirming size-distance invariance.

Other Forms of Perceptual Constancy

TypeWhat stays constantDespite changes in
Shape constancyPerceived shape of an objectViewing angle (a door swinging open still looks rectangular)
Colour constancyPerceived colour of surfacesIllumination (a red apple looks red under sunlight and artificial light)
Brightness constancyPerceived lightness of surfacesOverall illumination level (a white page looks white even in dim light)
Location constancyPerceived position of stationary objectsEye and head movements that shift the retinal image

Perceptual Adaptation

Perceptual adaptation is the ability of the perceptual system to recalibrate when it receives systematically distorted input. Stratton (1897) had participants wear goggles that inverted the visual field. Initially this caused disorientation. Over several days, participants adapted and their perception normalised. The key finding: active movement (physically interacting with the environment) drove adaptation, not passive exposure alone. The perceptual system recalibrates through sensorimotor feedback.

Section 06

Visual Illusions

IMP

Visual illusions are cases where the percept systematically diverges from physical properties. They are valuable because they reveal the assumptions the visual system uses. An illusion is a case where those assumptions fail.

Lab 05·A · Müller-Lyer
equal
Lab 05·B · Ponzo
equal
Lab 05·C · Ebbinghaus
equal
Fig.Three geometric illusions — reveal the true measurements
Müller-Lyer and Ponzo are constancy-scaling failures; Ebbinghaus is a contrast effect. Toggling the rules shows the physical equality your visual system refuses to accept.
IllusionPhenomenonExplanation
Müller-LyerTwo equal lines look different: outward arrowheads make a line look longerArrowheads create depth cues resembling near and far room corners. Size constancy scaling is applied to a flat figure, distorting length.
PonzoTwo identical bars between converging lines look different in sizeConverging lines provide a linear perspective depth cue. The upper bar appears farther, so constancy scaling makes it look larger.
Ebbinghaus (Titchener)A circle surrounded by small circles appears larger than the same circle surrounded by large onesContrast effect: the central circle is judged relative to surrounding context, not by absolute retinal size.
Moon IllusionThe moon looks larger near the horizon than high in the sky, despite identical angular sizeThe horizon moon appears farther (terrain depth cues intervene), so constancy scaling makes it look larger. The zenith moon lacks these cues.
Ames RoomPeople at opposite corners of a distorted room look radically different in sizeThe room fools the visual system into assuming a rectangular shape. When that assumption is wrong, constancy scaling produces large distortions.
Necker Cube / Rubin's VaseAmbiguous figures that spontaneously reverse between two interpretationsBoth interpretations fit the input equally well. The visual system generates competing hypotheses and alternates between them (Gregory).
Gregory's Misapplied Constancy (1963, 1966)

Many geometric illusions arise because the visual system misapplies size-constancy scaling to flat, two-dimensional figures. Cues in the figure (converging lines, arrowheads) trigger depth-processing mechanisms appropriate for 3D scenes; applied to a flat figure, they produce distortions. Gregory predicted that people with less exposure to carpentered (right-angle-rich) environments should show weaker illusions, a prediction with partial support (Segall, Campbell and Herskovits, 1966).

Induced Motion

A stationary object appears to move when its surrounding frame moves. The visual system uses the frame as a reference point and assigns relative motion to the smaller enclosed object. Classic example: when the train beside you moves, your stationary train appears to move in the opposite direction.

Helson's Adaptation-Level Theory (1964)

Adaptation level (AL) is a weighted average of all stimuli of a given class that the observer has recently been exposed to. Stimuli above the AL appear intense or large; those below appear weak or small; those at the AL appear neutral. The AL shifts with exposure: after lifting heavy weights, moderate weights feel lighter. This context-dependence is the same mechanism behind contrast-based illusions, where a stimulus looks different depending on what surrounds it.

Section 07

Speech Perception

IMP

In normal conversation, speakers produce about 10–15 phonemes per second. Listeners parse this stream effortlessly, recognise words, distinguish speakers, and track meaning simultaneously. Several features of the speech signal make this difficult to explain.

There is no acoustic invariance: the same phoneme sounds different depending on surrounding sounds (coarticulation), the speaker's voice, rate, and accent. There are no clear word boundaries in the continuous stream, yet listeners hear discrete units. Context also provides powerful top-down constraints: when one phoneme is replaced by a cough, listeners hear the complete word without noticing the gap. Warren (1970) called this the Phoneme Restoration Effect.

Theories of Speech Perception

TheoryMain ClaimMechanismStrengthLimitation
Motor Theory (Liberman et al., 1967)Listeners recover the speaker's intended articulatory gestures, not acoustic patternsA speech-specific module uses knowledge of one's own articulation to decode what the speaker intendedExplains categorical perception and the McGurk effectInfants and people with severe motor disorders still perceive speech normally
TRACE Model (McClelland & Elman, 1986)Three interactive levels: features, phonemes, and words. Activation flows both upward and downwardWord-level nodes feed back to activate expected phonemes, which activate expected features. This two-way flow explains context effects naturallySimulates human data well; explains phoneme restoration and the Ganong EffectComputationally complex; timing of top-down feedback is debated
Cohort Model (Marslen-Wilson & Tyler, 1980)Initial phonemes activate a "cohort" of all matching words, which narrows as more input arrivesContext and word frequency prune the cohort further until one word remainsExplains why listeners often identify a word before it finishes (word-onset effect)If the first phoneme is misheard, the target word never enters the cohort (addressed in the 1987 revision)
Direct Realist Theory (Fowler, 1986)Listeners directly perceive the articulatory gestures that produced the sound, without a special moduleArticulatory gestures are real events in the world; the acoustic signal is structured by them, so listeners evolved to perceive these events directlyAvoids a special speech module; integrates with Gibson's general ecological approachVague about how listeners "directly" perceive gestures from sound waves
Section 08

Signal Detection Theory

V. IMP

Traditional psychophysics assumed a fixed absolute threshold: below it, stimuli are never detected; above it, they always are. Signal Detection Theory (SDT) replaced this. Detection depends on two independent factors: the observer's sensitivity (d-prime) and their response criterion (beta), which is how cautious or liberal they are about reporting a signal. The same sensitivity can produce very different detection rates depending on the criterion.

Sensitivity (d-prime)

Measures how far apart the signal-plus-noise and noise-alone distributions are. A higher d-prime means the observer can more reliably tell signal from noise. Sensitivity is a true perceptual measure, independent of motivation or strategy.

Criterion (beta)

The decision cut-off: how much evidence the observer requires before saying "yes." When missing a signal carries severe consequences, the observer lowers the criterion, accepting more false alarms to avoid misses. Criterion is motivational, not perceptual.

Lab 06 · Signal detection explorerdrag the criterion line
NOISESIGNAL + NOISEd′ = 2.0CRITERION
d′ (sensitivity)"yes" to everything right of the line
Hit
84%
Miss
16%
False alarm
16%
Correct rejection
84%
Slide d′ and the distributions move apart: that is perception. Drag the criterion and all four percentages trade off with no change in sensitivity: that is motivation. SDT's whole point is that these are independent.
Fig.Move the criterion, watch the four outcomes trade off
Sliding the criterion changes hits and false alarms together while d-prime stays fixed. Separating the two curves changes d-prime, and only then does actual accuracy improve.

The Four Outcomes

Signal PresentSignal Absent
Respond "Yes"HitFalse Alarm
Respond "No"MissCorrect Rejection

SDT in practice

A radiologist reading chest X-rays knows that missing a tumour is catastrophic, so they lower the criterion and report any hint of a shadow. A security guard faces the opposite trade-off (false alarms waste searches) and raises the criterion. Both may have identical d-prime but different hit rates because their criteria differ. SDT separates this motivational factor from true perceptual sensitivity.


Section 09

High-Frequency Topics

Gestalt Laws (Proximity, Similarity, Closure, Common Fate)
Figure-Ground Distinction (Rubin's Vase)
Biederman's RBC / Geons
Müller-Lyer and Gregory's Misapplied Constancy
Monocular vs Binocular Depth Cues
Size-Distance Invariance / Emmert's Law
Signal Detection Theory (d-prime, beta, Hit/Miss/FA/CR)
Marr's Three-Stage Computational Model
Top-Down vs Bottom-Up Processing
Categorical Perception / McGurk Effect
Feature Detection Theory (Hubel and Wiesel)
TRACE Model / Cohort Model
Spatial Frequency Channels (Campbell and Robson)
Subliminal Perception and Priming
Associative vs Apperceptive Agnosia
Section 10

Previous Year Questions

UGC NET · 20191 / 7

Phi-Phenomenon is best seen between which of the following time intervals?

UGC NET

Phi-Phenomenon is best seen between which of the following time intervals?

UGC NET

Among the laws of perceptual grouping, the law of simplicity is a tendency to:

UGC NET

A man judged to be six feet tall when standing at ten feet away has a retinal image of size X. At twenty feet, the retinal image is X/2. How tall shall he be perceived at a distance of five feet?

UGC NET

Match the following: List I (Name): a. Template Matching Model b. Feature Matching Model c. Recognition by Components Model d. Configuration Model; List II (Feature): i. Spatial relations deviate from prototype ii. 3D objects described via parts and spatial relations iii. Visual analysis detects colours and edges iv. Recognition of barcodes.

UGC NET

Signal detection depends upon:

UGC NET

An important factor which enables one to adapt to inverted vision is:

UGC NET

Some people are able to draw an object, match similar objects and describe its component parts, but fail to recognize what they have just seen or drawn. This describes:

Section 11

Rapid Revision

Quick Revision
Tap any row to reveal the answer
Sensation, Perception & Gestalt
Gibson vs Gregory
Gibson: Direct/Ecological (bottom-up, no computation). Gregory: Constructivist (top-down, perception as hypothesis)
Prägnanz
Overarching Gestalt law: perceptual system organizes stimuli into the simplest, most stable form possible
Figure vs Ground
Figure: definite shape, lies in front, smaller/symmetric/convex. Ground: shapeless, extends behind, contour belongs to figure
Common Fate
Elements moving in same direction and speed are grouped; critical for motion-based figure-ground
Phi phenomenon optimal ISI
30–200 ms; below 30 ms: fusion; above 200 ms: sequential flashes
Rubin (1915)
Figure-ground and reversible figures; Rubin's Vase
Pattern Recognition
Template theory fails because
Needs enormous store; cannot recognize novel instances; rotation/scaling should impair recognition but doesn't
Hubel and Wiesel
Simple, complex, hypercomplex cells in V1; feature-detecting neurons; Nobel Prize
Pandemonium (Selfridge, 1959)
Four demon levels: image, feature, cognitive, decision; anticipated convolutional neural networks
Campbell and Robson (1968)
Multiple spatial frequency channels; selective adaptation: fatigue at one frequency only
Marr: 3 representations
Primal Sketch (edges/blobs), then 2½-D Sketch (viewer-centred depth), then 3-D Model (viewpoint-independent)
Biederman's RBC
~36 geons; objects = geon assemblies; viewpoint-invariant recognition; junctions critical for parsing
Depth, Constancy & Adaptation
Monocular depth cues
Linear perspective, texture gradient, occlusion, relative size, height in field, aerial perspective, motion parallax, accommodation
Binocular disparity
Difference between left/right eye images; basis of stereopsis; greatest for nearby objects
Emmert's Law
Perceived size of afterimage is proportional to perceived distance of surface; retinal size constant
Induced motion
Stationary object appears to move when surrounding frame moves; motion attributed to enclosed object
Stratton (1897)
Inverting goggles: active movement (not passive exposure) is the key factor in perceptual adaptation
Apperceptive vs Associative Agnosia
Apperceptive: perceptual stage fails (can't copy/match). Associative: percept intact, semantic disconnection (can draw but can't identify)
Illusions, SDT & Speech Perception
Müller-Lyer (Gregory)
Misapplied size constancy: arrowheads trigger room-corner depth cues; cultural variation (Segall, 1966)
SDT: sensitivity vs criterion
d-prime: sensitivity (actual ability to detect signal); beta: response criterion (motivational, decisional)
SDT 4 outcomes
Hit, Miss, False Alarm, Correct Rejection; performance = hit rate and false alarm rate together
Motor Theory (Liberman)
Listeners recover articulatory gestures, not sounds; explains categorical perception and McGurk effect
TRACE Model
Interactive activation: features, phonemes, words, with top-down feedback; explains phoneme restoration
Cohort Model (Marslen-Wilson)
Initial phonemes activate a candidate cohort; narrows until one word remains; explains early word recognition

Section 12

Theorist Quick Reference

Theorist(s)YearContributionKey Term
Wertheimer1912, 1923Founded Gestalt psychology; phi phenomenon; grouping lawsApparent Motion, Gestalt Principles
Rubin1915Figure-ground perception and reversible figuresRubin's Vase, Figure-Ground
Gibson, J.J.1950, 1979Ecological approach; optic array; affordances; texture gradientsDirect Perception, Affordances
Gregory1963, 1966Constructivist theory; misapplied constancy explanation of illusionsPerception as Hypothesis, Misapplied Constancy
Hubel and Wiesel1959–1962Feature-detecting neurons in primary visual cortexSimple, Complex, Hypercomplex Cells
Selfridge1959Pandemonium model of pattern recognitionFeature Demons, Cognitive Demons
Campbell and Robson1968Multiple spatial frequency channels in visionSpatial Frequency, Channel Adaptation
Marr1982Computational theory of vision; three-stage representationPrimal Sketch, 2½-D Sketch, 3-D Model
Biederman1987Recognition by Components (RBC)Geons, Viewpoint Invariance
Emmert1881Size-distance invariance; afterimage size scalingEmmert's Law
Brunswick1956Probabilistic functionalism; lens model of perceptionEcological Validity, Cue Utilisation
Helson1964Adaptation-level theoryAdaptation Level (AL)
Liberman et al.1967Motor theory of speech perceptionArticulatory Gestures, Categorical Perception
McClelland and Elman1986TRACE model of speech perceptionInteractive Activation, Top-Down Feedback
Marslen-Wilson and Tyler1980Cohort model of spoken word recognitionCohort, Recognition Point
Fowler1986Direct Realist Theory of speech perceptionArticulatory Events as Distal Objects
McGurk and MacDonald1976McGurk Effect: multimodal speech perceptionVisual-Auditory Integration
Warren1970Phoneme Restoration EffectTop-Down Lexical Completion
Segall, Campbell and Herskovits1966Cultural variation in susceptibility to geometric illusionsCarpentered World Hypothesis
Stratton1897Perceptual adaptation to inverted visionActive Movement, Perceptual Recalibration
Bruner and Minturn1955Perceptual set ("B vs 13" study)Perceptual Set, Context Effects
Perception & Attention — Free Psychology Note | The Exam Brief