This page examines fundamental problems that arise when spatial reality is viewed, imaged, projected, represented and interpreted. These include the relationships between physical space and image space, viewpoint and appearance, retinal images and visual experience, three-dimensional objects and two-dimensional representations, and real scenes and their possible reconstructions.
Many of these problems cannot be resolved by the rules of linear perspective alone. They concern the limits of projection, optical information, visual perception, instruments and representation, and remain important across art, science, photography, cinema, computer graphics, virtual reality and artificial intelligence.
Many misunderstandings of perspective do not arise from incorrect terminology alone. They arise from deeper problems concerning what a view or image can reveal, how spatial reality is transformed by projection, and how much information is lost, simplified or altered during seeing, imaging and representation.
Perspective never presents spatial reality in a completely neutral or exhaustive form. Every perspective view depends upon a viewpoint, viewing direction, projection system, image surface, scale, resolution and moment in time. A view may therefore be accurate within its stated conditions while remaining partial, ambiguous or misleading when those conditions are ignored. The following are not simply “errors” in perspective; they are fundamental problems that any adequate theory or interpretation of perspective must recognise.
The Equivalence / Correspondence Problem
The Equivalence or Correspondence Problem concerns the relationship between a spatial object or scene and the view or image produced from it.
A single monocular two-dimensional image of a three-dimensional object does not normally contain enough information to identify every feature of its source unambiguously. Different objects, shapes, dimensions, orientations or spatial arrangements can produce the same—or a sufficiently similar—projected image when viewed from particular positions.
A square, rectangle, trapezium or irregular plane may, under different viewing conditions, project to an identical quadrilateral. Likewise, an object’s apparent height may result from its physical size, its distance, its inclination or some combination of these factors. The image alone does not necessarily tell us which explanation is correct.
The central problem can therefore be stated as:
One image may correspond to several possible spatial realities.
Perspective provides information about the source, but it does not always provide a unique solution. To interpret the image, we employ context, prior knowledge and additional depth cues. Known dimensions, metric grids, ground planes, axes, horizons, vanishing structures, shading, occlusion and multiple viewpoints can all reduce uncertainty. However, there is no universal method by which the complete geometry of every three-dimensional object can be recovered unambiguously from one monocular image.
The terms equivalence and correspondence describe closely related aspects of the same problem. Equivalence concerns the fact that different spatial sources may produce equivalent image forms. Correspondence concerns the difficulty of deciding which image features correspond to which objects, positions and structures in the spatial source.
This problem is central to human vision, photography, graphical perspective, computer vision, photogrammetry, robotics, artificial intelligence and three-dimensional reconstruction.
Primary Spatial Factors
A perspective view initially requires us to identify at least three principal spatial factors:
- the distance or depth location of an object;
- its height or dimensional size;
- and its lateral position within the field.
Other factors include the object’s shape, orientation, angle, scale, motion and relationship to surrounding structures. Yet even these apparently elementary features may be difficult to determine from a single image because they interact.
A small image may represent a physically small nearby object or a much larger distant object. A shortened image may indicate a short object, an inclined object or a long object seen in foreshortening. Perspective interpretation is therefore a process of resolving interdependent spatial variables rather than simply reading fixed information directly from an image.
The Planes / Visual Angles Problem
Human vision and graphical linear perspective do not measure spatial appearance in precisely the same way.
The eye responds primarily to visual angle: the angular extent subtended by an object at the viewpoint. Its optical image is formed upon a curved retinal surface. Conventional linear perspective, by contrast, usually projects visual rays onto a flat picture plane and measures the resulting image using linear dimensions.
This creates the Planes / Visual Angles Problem:
The eye organises appearance angularly upon a curved retinal surface, while linear perspective commonly represents appearance through linear projection upon a flat plane.
A flat picture plane and a curved retinal surface cannot produce an identical mapping across an unlimited field. Near the central viewing direction, rectilinear projection can provide a close and highly useful correspondence. Across wider fields, however, differences become increasingly important. Peripheral directions, planar orientation, viewing distance, picture-plane location and retinal curvature all influence the resulting image.
This does not make linear perspective invalid. It means that its validity is conditional. A flat rectilinear projection accurately models particular geometrical relationships under specified viewing arrangements; it should not automatically be treated as a complete model of the entire visual field or of conscious visual experience.
The problem also explains why photographic, retinal, spherical, cylindrical and graphical images may differ even when they represent the same scene from approximately the same viewpoint.
The Problem of Space
Space itself is invisible. We do not see empty spatial extension directly; we infer it from the relationships between visible objects, surfaces, boundaries, changes and movements.
To make space understandable, measurable or reconstructable, perspective systems impose or identify spatial frameworks. These may include:
- axes and coordinates;
- horizontal and vertical planes;
- metric grids and checkerboards;
- lines of recession;
- vanishing points and vanishing lines;
- horizons;
- known dimensions and proportional relationships.
Such structures segment, order, measure and index spatial reality. They allow a perspective image to be encoded during construction and decoded during interpretation. Linear perspective is especially effective where an object space contains known framework structures, such as architectural edges, rectilinear surfaces or a ground-plane grid.
However, not every scene contains obvious axes, planes or regular objects. Natural environments may be irregular, curved, fragmented, partly hidden or lacking any known measure. In those cases, the spatial structure must be inferred through incomplete evidence.
The Problem of Space therefore concerns the gap between:
invisible spatial organisation
and
the visible information through which that organisation is inferred.
The dictionary further distinguishes physical, mathematical, artificial and imaginary spatial realities. A perspective image may concern a physical scene, a mathematical model, a human-made representation or an imagined space. Similar visual structures may occur in each, making the underlying nature of the represented space difficult to determine from appearance alone.
The Problem of Viewpoint
Every viewpoint possesses a particular optical and geometrical relationship to the scene. When the viewpoint changes, the image normally changes with it.
Object shapes become foreshortened, expanded, hidden or revealed. Relative positions alter. Occlusions change. Surfaces visible from one direction disappear from another. A single view consequently contains unique but partial information about its source.
This is the Problem of Viewpoint:
Every perspective view reveals some spatial information while excluding or transforming other information.
No individual viewpoint can normally show every side, surface, cavity or spatial relation belonging to a complex object. To understand the object sufficiently, we may require several views from different positions and directions.
Multi-view perspective can provide a fuller account, but it introduces further problems. The separate views must be matched, registered and integrated. We must determine which features correspond across the images and how their changing projected shapes relate to one stable three-dimensional structure. This returns us directly to the Equivalence / Correspondence Problem.
The Problem of Viewpoint also challenges the common assumption that one fixed station point provides a complete or uniquely correct account of visual reality. Fixed viewpoint is essential to many graphical and optical methods, but normal human vision involves eye, head and body movement, together with the accumulation of information across time.
The Problem of Time
Perspective is commonly discussed as though it concerns space alone. Yet every act of viewing, imaging or representation also occurs in time.
A drawing or photograph fixes one selected moment. A film records a sequence of moments at a defined frame rate. A scientific instrument may capture events lasting fractions of a second, while astronomy may examine processes extending across centuries or millions of years.
The Problem of Time concerns the narrow temporal window through which most perspective images represent changing reality.
An ordinary static image suppresses temporal development. A moving image restores movement, but it is still limited by:
- its frame rate;
- exposure time;
- recording duration;
- playback speed;
- selected starting and ending points;
- and the temporal scale of the process being observed.
Very rapid and very slow processes may therefore remain invisible or incomprehensible at ordinary viewing speeds.
Volume 1 proposes the need for multi-time perspective: systems that permit recorded time-flows to be indexed, linked, accelerated, slowed, combined and explored across different temporal scales. The viewer might move between nanoseconds, seconds, years or geological periods in a manner analogous to zooming through spatial scales.
The underlying principle is:
Every perspective image is not only a view from somewhere; it is also a view from some moment and across some duration.
Ignoring this temporal condition can produce a misleading impression that the represented scene is permanent, simultaneous or complete.
The Real / Simulated Problem
Modern visual systems increasingly combine physical, photographed, modelled, represented and computer-generated realities. A convincing perspective image may originate from a physical scene, a virtual model, an artificial construction or a combination of all three.
The Real / Simulated Problem concerns our ability to determine which parts of a view or image derive from physical reality and which have been simulated, altered or generated.
A photograph may be digitally reconstructed. A physical actor may appear inside a virtual environment. A real architectural space may incorporate projected scenery. A synthetic image may imitate the optical properties of a camera so convincingly that appearance alone no longer discloses its origin.
The difficulty is not merely distinguishing “true” from “false”. A simulated image can accurately model a real spatial process, while a photograph of a physical scene can be misleading because of viewpoint, lens choice, scale, cropping or context.
The more precise question is:
What kind of spatial reality produced this image, and through which sequence of physical, optical, graphical or computational processes?
The source may involve physical reality, mathematical reality, artificial reality or imaginary reality. Composite and synthetic systems may join several of these into one coherent appearance. Provenance, metadata, contextual evidence and knowledge of the image-forming chain are therefore increasingly necessary when appearance alone is insufficient.
The Scale / Shape / Size Problem
The Scale / Shape / Size Problem concerns the fact that perceived or measured shape and size cannot always be separated from the scale and resolution at which an object is viewed, imaged or measured.
The inverse size–distance law remains fundamental: under fixed projection conditions, the projected size of a specified dimension decreases as its distance increases. Doubling the distance halves the projected extent; trebling the distance reduces it to one-third.
However, this law does not by itself provide a complete account of the measured or apparent size of an irregular form.
An irregular coastline, stone or biological outline reveals progressively finer detail as the projection or measurement scale increases. A coarse measurement ignores small bays, cracks and deviations; a finer measurement follows more of them and may produce a longer boundary. The visible or measured shape therefore changes with the level of detail resolved.
The important principle is:
Measured or apparent size depends not only upon distance, but also upon projection scale, magnification and resolution.
A pebble viewed by the unaided eye may possess a relatively smooth outline. Under a microscope, the same boundary becomes jagged and structurally complex. The object has not simply “become larger”; a different level of form has become visible.
This is described as a problem of Extrinsic Simplicity because the apparent simplicity or complexity of the form depends partly upon external viewing, imaging and measuring conditions. The image of the object is made simpler or more complex by distance, scale, magnification and the resolving power of the system.
A complete perspective account must therefore state the scale and resolution at which size and shape have been defined.
Shape Sufficiency and Levels of Abstraction
Physical reality contains irregularities and complexities at many dimensional scales. A line that appears straight at architectural scale may be rough or fractured microscopically. A ground plane that is sufficiently flat for a perspective construction is not perfectly flat at every scale. Objects represented as simple solids may possess immensely complex surfaces, materials and internal structures.
Perspective systems manage this complexity through Shape or Form Sufficiency.
A geometrical form is sufficient when it captures the degree of structure required for a particular purpose and scale of analysis. Linear perspective may treat:
- lines as sufficiently straight;
- planes as sufficiently flat;
- parallel lines as sufficiently parallel;
- regular shapes as sufficiently regular;
- and surfaces as sufficiently continuous.
These are not claims that physical reality is perfectly Euclidean or geometrically simple. They are controlled abstractions that allow the relevant relationships to be represented, measured or analysed.
This produces different Levels of Abstraction. The same object may be represented as:
- a point indicating position;
- a line indicating direction or extent;
- an outline indicating shape;
- a plane indicating surface;
- a solid indicating volume;
- or a complex multi-scale model containing surface, material and internal detail.
Each representation omits information. Its validity depends upon whether the retained geometry is sufficient for the task.
This is described as Intrinsic Simplicity because simplification is built into the represented form or model itself. By contrast, Extrinsic Simplicity arises from external limits such as viewing distance, magnification and resolution.
Volume 1 argues that geometrical simplification underpins natural, artificial and synthetic perspective systems. Perspective succeeds not because it reproduces every detail of physical reality, but because its selected forms sufficiently represent the relevant order and relationships at a defined scale.
How the Problems Interconnect
These problems cannot be treated in isolation.
Changing viewpoint alters projected shape and may create new equivalences between different objects. Changing scale alters both visible detail and measured size. Changing the projection surface alters the relationship between visual angle and image dimensions. Time reveals transformations that remain hidden in a static view. Simulated systems can reproduce the visible outcome of physical processes without sharing their source. Every attempt to structure invisible space depends upon geometrical assumptions and selected levels of abstraction.
Consequently, no perspective image should be interpreted without considering:
viewpoint, direction, spatial framework, projection surface, scale, resolution, temporal interval, source reality and degree of abstraction.
A perspective view may be geometrically correct and still incomplete. It may be optically convincing and still ambiguous. It may be highly detailed and still fail to reveal its source. These are not exceptional failures; they are fundamental conditions of vision, imaging and representation.
Conclusion
Perspective is often presented as a transparent method by which three-dimensional reality is simply transferred into a two-dimensional image. In fact, every such transfer involves selection, transformation, information loss and interpretation.
The fundamental problems described here explain why perspective remains difficult even when its terminology is clearly defined. A view never merely shows an object “as it is”. It shows that object:
from a particular viewpoint, in a particular direction, at a particular scale and resolution, during a particular interval of time, through a particular optical or representational system.
Recognising these conditions does not weaken perspective. It establishes the limits within which perspective becomes rigorous, intelligible and useful.