Event cameras
Image sensors whose pixels each report a change in log intensity the moment it happens, with a microsecond timestamp — the retina's transient pathway in silicon, in place of frames.
A video camera reports the brightness of every pixel at every frame, whether or not anything has changed, and nothing about what happened in between. An event camera turns that around. Each pixel watches its own light continuously and speaks only when the logarithm of that light has moved by a set step since it last spoke, sending its address, the sign of the change and a microsecond timestamp. The first practical one, the dynamic vision sensor (DVS) of Patrick Lichtsteiner, Christoph Posch and Tobi Delbrück, was presented in 2006 and described in full in 2008. Its ancestor is the silicon retina.
A pixel that compares the light with its own past
A photoreceptor holds the photodiode at virtual ground through a feedback transistor running below threshold, so its output voltage follows the logarithm of the photocurrent, plus an offset that varies widely between pixels. A switched-capacitor amplifier amplifies only the change since its last reset, by a capacitor ratio of 20, and two comparators watch the result, one for brighter (ON) and one for darker (OFF). When either trips, the pixel puts its address on a shared bus — the address-event representation of neuromorphic chips — and the amplifier is reset, storing the present level as the new reference.
In the idealised model, a pixel fires as soon as
where is its photocurrent, the time of its previous event and the contrast threshold, typically 10–50%. It is a delta modulator on log intensity with no clock: the signal picks the sampling instants. Light at a pixel is illumination times reflectance, so under steady lighting an event marks a change of reflectance, usually a moving edge, and an edge of given contrast crosses as many thresholds in dim light as in bright. The 2008 chip had 128 × 128 pixels, each 40 µm square, worked over more than 120 dB — six decades of illumination — responded in as little as 15 µs in bright light, and drew 23 mW.
One retinal pathway
In 1971 Cleland, Dubin and Levick sorted the cat’s retinal ganglion cells into sustained and transient types, each with on-centre and off-centre versions, and found the two kinds kept essentially separate as far as visual cortex. The DVS models the transient kind, abstracting the path from photoreceptor through bipolar cell to ganglion cell and stopping there. Three habits crossed: reporting relative change rather than level; splitting brightening and darkening onto separate ON and OFF outputs, which saves the retina the firing that a single channel would need to signal both; and sending sparse spikes whose timing carries the message. What crossed was one pathway, as far as its ON and OFF ganglion cells.
Everything else stayed behind. The pixels ignore their neighbours, so there is no surround — none of the lateral inhibition that horizontal cells supply, and that Mahowald and Mead’s 1988 retina built from a resistive network. The mouse retina has more than 30 types of ganglion cell, each sending the brain its own version of the image (Baden and colleagues, 2016); the DVS sends one.
That economy was the point. Zaghloul and Boahen’s silicon retina reproduced the four major ganglion-cell types that drive visual cortex, but its pixel-to-pixel mismatch, a 2014 review judged, made it hard to use. Comparing each pixel with its own past cancels the photoreceptor’s offset, leaving mainly the comparators’ offsets, divided by the gain ahead of them; across the 2008 array the event threshold varied by 2.1% contrast.
What the events buy
The saving is in what is not sent. In the 2014 DAVIS paper, a 240 × 180 sensor watching a tennis backhand produced 60,000 events a second on average, peaking near 250,000, at a little over two bytes each: about 150 kB/s. A frame camera of the same resolution at 1,000 frames a second would send 54 MB/s, some 360 times as much, and still resolve time only to the millisecond, where events carry microsecond timestamps.
The timing is a temporal code: between two events at one pixel the log intensity moved by one threshold, so the interval measures how fast the light is changing.
Power follows the scene, up to a point. The analog front ends run continuously; what scales with activity is traffic. The DAVIS drew 5–14 mW depending on how much moved, most of it on the digital supply, chiefly the output pads, which took 1.2 mW at low activity and 8.3 mW at high. This is efficient coding done per pixel and without a clock: send the change, not the level the receiver already has.
No picture without extra circuitry
A DVS never reports a level, so a still camera facing a still scene sends only noise events. The retina shares the blind spot: an image held still on it fades rapidly, as Ditchburn and Ginsborg (1952) and, independently, Riggs and colleagues (1953) found, and ordinary eye movements reduce the fading. In 2024 He and colleagues copied the remedy, spinning a wedge prism in front of an event camera so that a still scene keeps producing events.
Chip designers added circuits instead. In the ATIS of Posch, Matolin and Wohlgenannt (2011) a change event resets a capacitor in the pixel, a second photodiode discharges it, and two more events mark its crossing of two thresholds, the interval between them inversely proportional to the light: the integrate-and-fire arithmetic, run once per change. Only changed pixels are re-measured, which compressed static scenes up to a thousandfold, with 143 dB of dynamic range, at the price of twice the pixel area and three times the event data. The DAVIS of Brandli and colleagues (2014) shares the photodiode with a conventional active-pixel readout, adding about 5% to the pixel for ordinary frames beside the events, with 51 dB of dynamic range against the events’ 130.
What it cost
The first cost was silicon. A 40 µm pixel with a 9.4% fill factor throws away about nine photons in ten, where conventional industrial pixels are 2–4 µm. Stacking fixed most of that: a Sony–Prophesee design shown in 2020 reached 4.86 µm with a fill factor above 77%, and in September 2021 Sony announced the 1280 × 720 IMX636 and the smaller IMX637, made with Prophesee, each pixel’s change detector on a logic chip beneath the pixel chip, joined by copper-to-copper bonds.
The output is sparse only while the scene is. Event rate rises with motion and texture, and a saturated output bus delays events and perturbs their timestamps; hence the 2020 design’s readout of 1.066 billion events a second and, in the commercial parts, an event-rate control that trims the stream to what downstream systems can process.
Noise is worst at the dim end of the range. Junction leakage drifts pixels into spurious ON events at a rate that climbs with temperature (Nozaki and Delbrück, 2017), and in dim light the photocurrent’s shot noise is relatively large. Sony specifies the IMX636 at 1 klux and at 5 lux: a background rate of 0.1 and 10 events per pixel per second — across 921,600 pixels, about 90,000 and 9 million events a second with nothing moving — and a latency under 100 µs and under 1 ms.
Nor is the output an image. An edge moving along its own length produces no events, polarity depends on the direction of motion, and frame-based algorithms expect dense images at fixed times. Accumulating events into frames reuses those algorithms but quantises the timestamps and can discard the sparsity; filters that update an estimate with each event, as a Kalman filter does with each measurement, keep the latency low but need a model of the scene to compare against, and of the noise, which the 2022 survey found nobody could yet predict under arbitrary lighting.
Origins & further reading
- Patrick Lichtsteiner et al., 2008. A 128 × 128 120 dB 15 µs Latency Asynchronous Temporal Contrast Vision Sensor. IEEE Journal of Solid-State Circuits. paper · doi
- P. Lichtsteiner et al., 2006. A 128 × 128 120 dB 30 mW asynchronous vision sensor that responds to relative intensity change. IEEE International Solid-State Circuits Conference, Digest of Technical Papers. paper · doi
- B. G. Cleland et al., 1971. Sustained and transient neurones in the cat's retina and lateral geniculate nucleus. The Journal of Physiology. paper · doi
- Christoph Posch et al., 2011. A QVGA 143 dB Dynamic Range Frame-Free PWM Image Sensor With Lossless Pixel-Level Video Compression and Time-Domain CDS. IEEE Journal of Solid-State Circuits. paper · doi
- Christian Brandli et al., 2014. A 240 × 180 130 dB 3 µs Latency Global Shutter Spatiotemporal Vision Sensor. IEEE Journal of Solid-State Circuits. paper · doi
- Thomas Finateu et al., 2020. A 1280×720 Back-Illuminated Stacked Temporal Contrast Event-Based Vision Sensor with 4.86µm Pixels, 1.066GEPS Readout, Programmable Event-Rate Controller and Compressive Data-Formatting Pipeline. IEEE International Solid-State Circuits Conference (ISSCC). paper · doi
- Kareem A. Zaghloul & Kwabena Boahen, 2006. A silicon retina that reproduces signals in the optic nerve. Journal of Neural Engineering. paper · doi
- Tom Baden et al., 2016. The functional diversity of retinal ganglion cells in the mouse. Nature. paper · doi
- R. W. Ditchburn & B. L. Ginsborg, 1952. Vision with a Stabilized Retinal Image. Nature. paper · doi
- Lorrin A. Riggs et al., 1953. The Disappearance of Steadily Fixated Visual Test Objects. Journal of the Optical Society of America. paper · doi
- Botao He et al., 2024. Microsaccade-inspired event camera for robotics. Science Robotics. paper · doi
- Yuji Nozaki & Tobi Delbrück, 2017. Temperature and Parasitic Photocurrent Effects in Dynamic Vision Sensors. IEEE Transactions on Electron Devices. paper · doi
- 2021. Sony to Release Two Types of Stacked Event-Based Vision Sensors with the Industry’s Smallest 4.86μm Pixel Size for Detecting Subject Changes Only. Sony Semiconductor Solutions news release. web
- Christoph Posch et al., 2014. Retinomorphic Event-Based Vision Sensors: Bioinspired Cameras With Spiking Output. Proceedings of the IEEE. paper · doi
- Guillermo Gallego et al., 2022. Event-Based Vision: A Survey. IEEE Transactions on Pattern Analysis and Machine Intelligence. paper · doi
Concepts
Related