AI Perspective

AI Perspective, or Artificial Intelligence Perspective, concerns the use of artificial intelligence systems to generate, analyse, identify, transform, connect, validate or otherwise work with perspective views, images and spatial models.

Artificial intelligence can now participate in both major directions of perspective. It can analyse existing views and images in order to obtain information about spatial reality, and it can generate new artificial perspective images, viewpoints, moving images and simulated spatial environments.

AI Perspective therefore extends perspective beyond the traditional relationship between an observer, object and picture plane. A modern perspective system may include cameras, sensors, artificial intelligence, computer vision, graphical generation, three-dimensional models, displays and robotic action within one continuous perspective image chain.


What Is AI Perspective?

The Dictionary of Perspective defines Artificial Intelligence Perspective as the use of an Artificial Intelligence system to generate artificial or counterfeit perspective images, generate novel perspective viewpoints or movies, or analyse image features.

The term therefore covers several different operations that should not be confused with one another:

  • AI image generation — creating artificial perspective images.
  • Novel-view generation — producing new perspective viewpoints from existing information.
  • Image analysis — extracting information about objects, scenes and processes from perspective images.
  • Image identification and classification — determining what objects, forms or events are represented.
  • Image validation — assessing or certifying information contained within views and images.
  • Image integration — connecting large numbers of views into larger image spaces or models.
  • Robotic vision and action — using analysed perspective information to guide machines in physical space.

AI Perspective is consequently much broader than simply making pictures with generative AI.


Artificial Intelligence and the Wider Field of Perspective

Perspective concerns the ways in which spatial objects, scenes, views and images are observed, formed, projected, measured, represented, simulated, transformed and experienced.

Artificial intelligence can now participate in many of these operations. It can analyse captured images, classify their contents, infer spatial relationships, generate graphical imagery, combine views, construct models and connect visual information with language and physical action.

This makes AI one of the most significant recent developments within New Media Perspective: computational and electronic systems through which perspective is generated, processed, displayed or explored.


AI Perspective and Artificial Perspective

Artificial Perspective is the wider term for perspective produced, modified, measured, projected, simulated or represented through deliberately designed methods, systems, instruments or spatial arrangements.

Importantly, such systems need not be operated only by humans. They may be created or operated by humans, Artificial Intelligence systems, machines or combinations of these.

AI Perspective can therefore be understood as an important modern branch of artificial perspective in which artificial intelligence participates directly in creating, analysing or operating the perspective process.

The resulting image chain may nevertheless also involve Natural, Optical, Instrument, Mathematical, Graphical, Simulated and Visual Perspective. The presence of AI does not mean that every other part of the perspective system becomes artificial or belongs to one single category.


AI-Generated Perspective Images

One of the most visible applications of artificial intelligence is the generation of new images.

AI Drawing uses artificial intelligence technologies to create images, graphics or artwork. A system may generate a visual image from a textual description or create new graphical results based upon learned image information.

From the perspective viewpoint, the important point is that the resulting image may represent an apparent three-dimensional spatial reality even though no corresponding scene was directly viewed or photographed.

AI can therefore produce an artificial perspective image whose apparent viewpoint, spatial arrangement, objects, lighting and represented depth are generated computationally.


Counterfeit and Artificial Perspective Images

Volume 1 uses the expression counterfeit perspective view or image for an artificially generated perspective representation that can possess photographic, three-dimensional or moving-image qualities without being a conventional direct capture of the represented scene.

In this perspective sense, “counterfeit” identifies the artificial generation of the view rather than establishing any particular intention to deceive.

The distinction becomes increasingly important as artificial systems create images that may visually resemble ordinary photographs, films or rendered views while originating from completely different perspective processes.

Understanding such an image therefore requires asking not merely what does the image look like? but also how was this perspective image formed?


Generative AI and Perspective

Generative AI adds a new image-forming process to the field of perspective.

A traditional camera captures an optical perspective view of an existing target. A graphical artist constructs a represented view. A three-dimensional computer model generates an image according to defined digital geometry. Generative AI provides another possibility: a system can generate a new visual representation computationally from learnt information and supplied instructions.

The result can still exhibit familiar perspective phenomena including apparent depth, viewpoint, scale, shape, occlusion, spatial arrangement and forms of graphical convergence, but the route by which the image has been produced differs from ordinary optical capture.


Generating Novel Perspective Viewpoints

Another major application of AI is the generation of novel perspective viewpoints.

Instead of generating an unrelated new image, a system may begin with one or more existing views and produce another viewpoint of the same or corresponding spatial scene.

This changes the role of perspective images substantially. A set of captured views can become source information from which further views are generated.

The process can be represented as:

Existing Perspective View(s) → AI / Computational Analysis → Spatial Representation → Novel Perspective View.

Volume 1 identifies novel-view synthesis as one of the important existing applications of artificial intelligence to perspective.


AI Perspective and Computer Vision

Computer Vision is one of the principal analytical forms of AI Perspective.

Computer vision enables computers to analyse visual data in order to identify and interpret objects, scenes and processes represented within images and videos.

From the perspective viewpoint, the system attempts to extract useful information from perspective images of spatial reality. This may include information about:

  • object and scene location;
  • size and shape;
  • orientation;
  • focus;
  • spatial arrangement;
  • texture;
  • colour;
  • whole and part;
  • figure and scale;
  • position;
  • state; and
  • activity.

Computer vision therefore turns the perspective image from a passive visual representation into a source of data for measurement, classification, modelling and decision-making.


AI Image Classification

Image classification is one of the basic operations through which AI interprets perspective imagery.

A system can analyse visible image properties and attempt to determine the type or class of object, scene or process represented.

This relates directly to the perspective function of matching: comparing, measuring, classifying and cross-matching visual information in order to obtain useful knowledge about a target.

Image classification may operate upon individual images or upon very large image databases.


Object Identification and Recognition

Artificial intelligence can also perform automatic object identification within still or moving perspective images.

Volume 1 distinguishes object identification used in Machine-Vision and Robotic-Vision systems from related identification procedures operating upon image databases.

The perspective problem is significant because one physical object can present many different image forms according to viewpoint, distance, direction, scale, occlusion and other conditions.

AI therefore attempts to associate changing perspective appearances with a corresponding object, scene or process.


AI and the Certification of Perspective Images

Another possible function of AI Perspective is certification or validation of perspective views and images.

The Dictionary of Perspective defines Certification of Vision as techniques intended to establish or validate the truth of information obtained from views or images of physical reality.

This may involve validating information concerning whole and part, figure, size, scale, position, state or activity.

Such certification may apply to images obtained through direct human vision, optical instruments, New Media systems or combinations of these.

As artificial and captured images become increasingly interconnected, identifying the process through which an image was generated, transformed and interpreted becomes an increasingly important part of perspective analysis.


AI and Integrated Perspective Image Spaces

Artificial intelligence can also help link large numbers of separate perspective images into a larger coordinated image space.

This is important because any single perspective image normally contains only a restricted view. Different photographs, camera positions, scales, times and sensing systems can provide additional information about the same spatial reality.

AI and other New Media methods can help connect, order, match, classify and navigate these images so that they function as parts of a larger spatial representation.

The result may move perspective away from the traditional isolated picture towards a networked, multi-view and potentially explorable image space.


AI and Multi-View Perspective

New Media and Multi-View Perspective connects multiple images of a spatial reality through processes such as linking, ordering, constructing, matching, mixing, exploring and cross-matching.

Artificial intelligence can assist many of these functions by analysing relationships between views and identifying corresponding image information.

Multiple perspective views can consequently become components of a larger model rather than remaining unrelated individual images.


AI Analysis and Reverse-Engineering of Spatial Geometry

Volume 1 identifies sophisticated analysis of perspective images and views as another important application of artificial intelligence.

Such analysis may contribute to the reverse-engineering of information about the spatial object or scene represented within an image.

The problem can be expressed as:

Spatial Reality → Perspective Image → AI Analysis → Estimated Information about Spatial Reality.

This is closely related to the wider Problem of Reality within perspective: how can information about a three-dimensional object or scene be recovered from images in which depth, scale, shape, position and visibility have already been transformed by perspective?


AI Perspective and 3D Reconstruction

AI Perspective can also contribute to 3D reconstruction by analysing perspective views and helping infer spatial relationships from them.

A set of views can provide information about changing object position, apparent shape, visibility and viewpoint. These relationships can be organised into a more comprehensive model of the spatial target.

The reconstructed model can then become the basis for new views, measurements and representations.

This creates a circular relationship:

Spatial Reality → Perspective Images → Analysis / Reconstruction → Spatial Model → New Perspective Images.


AI and Navigation through Perspective Images

Another function identified in Volume 1 is the efficient navigation of perspective images.

An AI system may assist in combining, linking, orienting and navigating visual information obtained from different images, viewpoints or image spaces.

This is particularly important when the amount of visual information exceeds what can conveniently be examined as a simple sequence of independent pictures.

The image collection can instead become a structured visual environment through which information is located and explored according to spatial or informational relationships.


AI and Contextual Visual Overlays

AI Perspective can also combine a perspective view with additional contextual instruments and graphical information.

Volume 1 gives examples including scales, meters, throttles and steering controls overlaid upon images.

Such systems combine the original perspective view with graphical or instrumental information intended to aid measurement, interpretation or action.

This is especially relevant to augmented displays, machine interfaces and other systems in which the image is not merely viewed but becomes part of a larger information environment.


AI and Levels of Visual Abstraction

Artificial intelligence can also assist in the navigation and visualisation of different levels of abstraction.

A spatial subject can be represented in many different ways: as a photograph, outline, diagram, technical drawing, map, depth representation, model or other graphical abstraction.

AI systems can potentially help connect and organise such different representations so that the same target is explored through several types of visual or informational description.

This extends perspective beyond the production of one particular image form towards the coordinated exploration of different representations of spatial reality.


AI Perspective and Robotic Vision

Robotic Perspective refers to a robot’s or other artificial system’s view of the world and its associated visual sensing and computer-vision system.

An autonomous vehicle, android or other robot may use cameras and sensors to obtain perspective information about its surroundings and then analyse that information computationally.

The key difference from an ordinary camera is that the resulting image can become part of a decision-and-action system.

A simplified chain is:

View → Identify → Locate → Interpret → Decide → Act.

Perspective therefore becomes directly connected with physical behaviour within spatial reality.


Vision-Language Models

Vision-Language Models, or VLMs, combine computer vision and natural-language processing so that visual information can be analysed in relation to language.

Such systems can perform tasks including image analysis, captioning, visual question answering and image generation.

From the perspective viewpoint, the important development is that information contained within a perspective image can now be connected directly with linguistic information.

The image can therefore become something that an artificial system not only analyses geometrically or visually but also describes, queries and relates to other knowledge.


Vision-Language-Action Models

Vision-Language-Action models, or VLAs, extend the relationship further by linking visual information and language with physical action.

Volume 1 gives the simple example of a robot receiving an instruction such as “pick up the bottle”. To execute the instruction successfully, the system must relate language to visual perception, identify the object, establish relevant spatial information and control an action.

A VLA can therefore combine:

  • a vision encoder for visual information;
  • a language model for linguistic interpretation; and
  • an action decoder for controlling physical behaviour.

This provides one of the clearest examples of perspective information becoming part of an integrated cycle of perception, understanding and action.


AI-Generated Training Worlds

AI Perspective also includes the generation of synthetic three-dimensional and visual training data for robots and autonomous systems.

Artificial environments can present machines with large numbers of spatial situations, viewpoints, objects and events. These simulated worlds can be used to train systems to recognise unusual or difficult situations that may be insufficiently represented within ordinary captured data.

The resulting perspective images are artificial, but they can be used to develop machine responses to events expected within physical spatial reality.

This establishes an important perspective chain:

Artificial Spatial World → Generated Perspective Images → AI Training → Robotic / Machine Vision → Physical-World Action.


Optical Artificial Intelligence

The Dictionary of Perspective also identifies Optical Artificial Intelligence.

One meaning concerns light-based computing technologies, including photonic systems used to increase the speed or efficiency of artificial-intelligence processing.

A second meaning concerns AI applied directly to advanced imaging, including:

  • generation of synthetic visual training data;
  • generation of video and moving images;
  • generation of real-time interactive virtual environments;
  • analysis of the visual and physical world;
  • human assistance through augmented systems; and
  • scientific geometrical analysis across different physical scales.

Optical AI therefore represents an important meeting point between Artificial Intelligence, Optical Perspective, Instrument Perspective and New Media Perspective.


AI Perspective and Virtual Environments

Artificial intelligence can contribute to the production of real-time interactive virtual worlds and other artificial spatial environments.

Such environments can combine generated geometry, imagery, viewpoint change, interaction and responsive spatial behaviour.

The viewer may therefore no longer receive one predetermined image. New perspective views can instead be generated as the virtual viewpoint changes.

This connects AI Perspective with Virtual, Augmented, Mixed and Extended Reality and with the broader development of interactive New Media Perspective.


AI Perspective and the Perspective Image Chain

AI is often only one stage within a much larger perspective system.

For example:

Physical Scene → Optical Camera Image → Digital Image → AI Analysis → Spatial Model / Classification → New Image or Action → Display → Human Visual Perception.

Alternatively, the front of the chain may itself be artificial:

Text / Data / Artificial Model → AI Generation → Artificial Perspective Image → Digital Display → Human Visual Perception.

Understanding an AI perspective image therefore requires identifying the different spaces, transformations and perspective methods through which it has passed.


AI Perspective and Perspective Functions

Artificial intelligence can contribute to many of the fundamental functions of perspective.

  • Capturing and observing — through AI-assisted camera and sensor systems.
  • Measuring and calculating — through computational analysis of perspective information.
  • Classifying — through image and object recognition.
  • Modelling — through reconstruction and integration of spatial information.
  • Indexing and linking — through organisation of large image collections.
  • Certifying — through validation of image-derived information.
  • Exploring — through interactive image spaces and models.
  • Displaying — through generated or transformed visual representations.
  • Projecting and representing — through artificial images, virtual views and graphical outputs.

The products of these functions may include a visual image, measurement, calculation, representation, model or view of a three-dimensional object or scene.


AI Perspective and Perspective Category Theory

Within Perspective Category Theory, AI Perspective should not be treated as an additional principal category alongside Natural, Visual, Optical, Mathematical, Graphical, Instrument, Simulated and New Media Perspective.

Artificial intelligence operates principally as a New Media Perspective process, but an AI-based perspective system can involve several categories at once.

  • Natural Perspective may provide the original physical scene.
  • Optical Perspective may form an image through light.
  • Instrument Perspective may operate through cameras, sensors or displays.
  • Mathematical Perspective may calculate spatial and geometrical relationships.
  • Graphical Perspective may provide generated or represented imagery.
  • Simulated Perspective may provide artificial objects, environments or spatial effects.
  • New Media Perspective provides the principal computational environment for AI processing and generation.
  • Visual Perspective Type 2 operates when the final result is viewed and perceived by a human observer.

AI Perspective therefore provides an especially clear example of category chaining: several perspective processes may operate sequentially within one complete system.


AI Perspective versus Computer Vision

AI Perspective and Computer Vision are not identical.

Computer vision is primarily concerned with analysing visual data and extracting information about objects, scenes and processes.

AI Perspective is broader. It includes computer vision but also encompasses artificial image generation, novel-view generation, image certification, visual integration, navigation through image spaces, robotic interaction and other AI-based perspective methods and systems.

Computer vision can therefore be understood as one major analytical component within the wider field of AI Perspective.


Ten Major Applications of AI Perspective

Volume 1 identifies ten important ways in which artificial intelligence may be applied to perspective:

  1. Generate artificial or counterfeit perspective views and images, including photographic-quality, three-dimensional and moving imagery.
  2. Generate novel perspective viewpoints from existing viewpoints.
  3. Certify or validate perspective images and views.
  4. Automatically identify objects for Machine-Vision and Robotic-Vision systems.
  5. Automatically identify and classify objects within image databases.
  6. Integrate large numbers of perspective views into a connected image space.
  7. Analyse perspective images and views, including reverse-engineering aspects of scene or object geometry.
  8. Navigate perspective images by combining, linking and orienting visual information.
  9. Overlay contextual instruments and information onto perspective imagery.
  10. Navigate and visualise different levels of abstraction.

These functions demonstrate why AI Perspective is significantly wider than the popular idea of AI as merely an image-generation tool.


AI Perspective — Frequently Asked Questions

What is AI Perspective?

AI Perspective is the use of Artificial Intelligence systems to generate, analyse, identify, transform, connect or validate perspective images, viewpoints, movies and spatial information.

Is AI Perspective the same as generative AI?

No. Generative AI is one important part of AI Perspective. AI Perspective also includes computer vision, object identification, robotic vision, novel-view synthesis, image validation, image integration and other analytical or spatial functions.

Is AI Perspective a principal Perspective Category?

No. Within Perspective Category Theory, Artificial Intelligence operates principally within New Media Perspective and can participate in systems that also involve Optical, Instrument, Mathematical, Graphical, Simulated, Natural and Visual Perspective.

Can AI create perspective images?

Yes. AI systems can generate artificial graphical images representing apparent three-dimensional scenes and can also generate moving imagery and virtual environments.

Can AI generate a new viewpoint?

Yes. Volume 1 identifies novel-view synthesis as an important AI perspective function in which new perspective viewpoints can be generated from existing view information.

How is computer vision related to AI Perspective?

Computer vision is a major analytical form of AI Perspective. It uses computers to analyse visual data and identify or interpret objects, scenes and processes represented within perspective images.

What is Robotic Perspective?

Robotic Perspective concerns the view of the world obtained through a robot’s or artificial system’s visual sensing and computer-vision system, such as that used by an autonomous vehicle or robot.

What is a Vision-Language Model?

A Vision-Language Model combines computer vision and natural-language processing so that visual images and videos can be analysed in relation to language.

What is a Vision-Language-Action Model?

A Vision-Language-Action model combines visual information, language and physical control so that an artificial system can interpret an instruction, understand the visual environment and perform an action.

Can AI validate perspective images?

Image certification and validation are identified as potential functions of AI Perspective. These processes concern establishing or testing information obtained from perspective views and images.

What is Optical Artificial Intelligence?

Optical Artificial Intelligence can refer either to light-based AI computing technologies or to the application of AI to advanced imaging, including synthetic visual training data, generated video, interactive virtual environments and analysis of the physical world.

Can AI Perspective help with 3D reconstruction?

Yes. AI analysis of perspective views can contribute to recovering or modelling spatial information, while reconstructed spatial models can subsequently be used to generate new perspective views.

Why is perspective important to AI?

Much of the visual information analysed or generated by AI consists of perspective views of spatial objects, scenes and processes. Perspective provides a framework for examining their viewpoint, direction, size, shape, scale, position, depth, spatial arrangement and relation to the reality or model from which they originate.


AI Perspective within the Wider Field of Perspective

Artificial Intelligence changes both sides of the perspective problem. Machines can increasingly interpret perspective images of spatial reality, while at the same time generating completely new perspective images, viewpoints, moving scenes and artificial environments.

AI can classify what an image contains, identify objects and processes, connect multiple views, reconstruct or analyse spatial relationships, generate new viewpoints, validate visual information and translate perspective knowledge into physical robotic action.

AI Perspective therefore represents a major expansion of the traditional field: from perspective as something principally seen, drawn, photographed or projected towards perspective as something that can also be generated, interpreted, connected, modelled and acted upon by artificial systems.