In psychology and neuroscience, multiple object tracking (MOT) refers to the ability of humans and other animals to monitor multiple moving objects. It is also the term for certain laboratory techniques used to study this ability. In an MOT study, several identical moving objects are presented on a display. Some of the objects are designated as targets while the rest serve as 'distractors'. The study participants try to monitor the changing positions of the targets as they and the distractors move about. At the end of the trial, typically the participants are asked to indicate the final positions of the targets. The results of MOT experiments have revealed limitations on humans' ability to monitor multiple moving objects simultaneously. For example, awareness of features such as color and shape is disrupted by the objects' movement.
Background
History In the 1970s, researcher Zenon Pylyshyn postulated the existence of a "primitive visual process" in the human brain capable of "indexing and tracking features or feature-clusters". Using this process, cognitive processes can continuously refer to, or "track", objects despite movement of the objects causing them to stimulate different visual neurons over time. Data collected with Pylyshyn's MOT protocol and published in 1988 provided the first formal demonstration that the mind can keep track of the changing positions of multiple moving objects. As a specific theory of this ability, Pylyshyn proposed "fingers of instantiation" theory (FINST), which is that tracking is mediated by a fixed set of discrete pointers. Whereas FINST theory has been very influential, many studies have found evidence that seems inconsistent with the theory.
Procedure
A typical MOT study involves the presentation of between eight and twelve objects. The participant is told to monitor the positions of a subset of the objects, which are referred to as targets. Often the targets are indicated by being presented initially in a distinct color. The targets then become identical in appearance to the other, distractor objects. The targets and distractors move about the screen for several seconds in an unpredictable fashion. The participant is then asked to indicate which of the objects are the targets. The accuracy of the participant's judgments indicates whether the participant mentally updated the positions of the targets as they moved. To ensure that the task requires participants to mentally update the targets' positions, displays are typically designed such that object paths cause the targets to swap positions with distractors, at least occasionally. With that constraint, MOT task variations have been designed to probe specific aspects of how the mind tracks moving objects. For example, to compare performance in the left to performance in the right visual fields, studies confine some or all the moving objects to one of the visual fields. To avoid any contribution from spatial interference among mental object representations, some studies maintain a minimum distance between objects. Other studies have combined MOT with a concurrent task to investigate whether the two tasks draw on the same mental resource, and have changed target features such as color to assess whether study participants update their representations of those features.
Capacity limits MOT study results indicate that the number of targets that people can track is very limited. This reflects a bottleneck in the brain's processing architecture. Whereas at the early, sensory stages of visual processing, dozens of objects may be fully processed, later processes such as those associated with cognition have much more limited capacity to process visual objects. The specific number of visual objects that people can accurately track varies widely with display parameters, contrary to a common belief that people can track no more than four or five objects. Even for a fixed set of display parameters, rather than there being a clear limit, performance falls gradually with the number of targets. Such findings undermine Pylyshyn's FINST theory that tracking is mediated by a fixed set of discrete pointers. The above limitations appear to stem from processes specific to the two cerebral hemispheres. The independence of the limits in the two hemifields is demonstrated by findings that when one is tracking the maximum number that can be tracked in the left hemifield (which is processed by the right cerebral hemisphere), one can add targets to the right hemifield (which is processed by the left cerebral hemisphere) at little to no cost to performance. For features other than position, capacity seems to be more limited—see § Updating of features other than position. Whereas the tracking capacity limit is largely set separately by the two cerebral hemispheres, a more unified and cognitive resource also can contribute to tracking. For example, if there is only one target, one can bring one's full cognitive abilities to bear, such as in predicting future positions, to facilitate tracking. When more targets are present, these resources may still play a role.
… excerpt ends here. Continue reading the full article.

