This is a "computer scientists" understanding of photography, and this phrase alone can even be seen as "dangerous" by photographers.
The actual headline of the article is actually more informative and less misleading than the subhead:
"Unbounded High Dynamic Range Photography Using a Modulo Camera"
In photographer's terms, they're referring to ending blown highlights (caused by filling the wells in a sensel) not saturated images (caused by exaggerating color data).
There's a school of photography that basically lives and dies by over-saturating images (e.g. http://kenrockwell.com ) and they'd likely have a cow over the headline, but understand and be totally onboard for the actual goal.
Depending on how quickly the well can reset and how many times resets can be counted this will either be better or worse than simply improving existing approaches. The state of the art in full frame sensors is just under 15-bits of dynamic range. So to equal it an 8-bit sensor would need to be able to count resets at least 128 times AND be able to do this in 1/8000 of a second, at the same sensel density and efficiency! Don't hold your breath.
(I'd also suggest that trends towards using multiple sensors to assemble high resolution images will easily blow this away since they can use high- and low- sensitivity sensors to synthesize more dynamic range and resolution. Also see recent Olympus patents to capture polarization data at the same time.)
BTW: assuming they can do this stuff, they presumably can read the sensor pretty darn fast — so they should be able to eliminate image "tearing" in digital video and the need for mechanical shutters altogether. That's probably a bigger issue than dynamic range.
I suspect that clearing a well in < 1/10,000,000 of a second (which would only lose 10% of the photon detection time during a 1/8000s exposure) is going to be tough. Assuming 25% of the sensel real estate needs to be sacrificed to the modulo circuitry, that's a loss of 32.5% of photon data which is about the same loss as for pellicle mirrors (as seen in Sony's pretty unsuccessful SLT cameras) for a feature that won't be much use in many situations.
There's a school of photography that basically lives and dies by over-saturating images (e.g. http://kenrockwell.com ) and they'd likely have a cow over the headline, but understand and be totally onboard for the actual goal.
Actually, I suspect that even that crowd, after the initial shock would be absolutely on board. This would mean that no longer would the sensor control how the high-end clipping happens, but now that could be tailored specifically via a post-processing filter. E.g. to exactly match a long-gone film emulsion's behavior, "enhance" a classic, or to simply to create new clipping functions.
This is a bit analogous to the flexibility that black and white shooters gain by working with modern color sensors. When shooting B&W film, the tonal result was controlled via a combination of the film's spectral response and any colored lens filters applied, the latter used to manipulate tonal contrast between different colored subject matter (e.g. sky, plants, skin tones, etc.). Now a skilled photographer can shoot color RAW w/ B&W preview (for a preview of the image luminosity) then adjust in post to nail the image effect without having to mess even thinking about what colored lens filter(s) to have used.
Good point -- if this worked even moderately well -- e.g. it lets you capture 20 bits of DR -- a photographer could go crazy with the old film approach of "expose to the right" and simply keep settings such that every shot is "over" exposed and simply fix it all in post (it's pretty hard to over-expose by more than 6 stops unless you're trying...)
To emulate film stock we need a solid 14-bits of Dynamic range (while DxO says the D810 has it, the article below says 11-bits, but the best video cameras are getting there):
Ideally we'd get something like this AND capture polarization data as well (per the recent Olympus patent). Then you can apply polarization and exposure correction in post.
I identified the same issue with cycle time, but to me the biggest problem is the assumption of an accurate, sensitive, non-destructive-read sensor. They're assuming that they can charge a capacitor from a photodiode accurately, and also hook a voltage comparator (which is effectively the same as an ADC) to it constantly to get an accurate pulse to count.
Assuming you have an accurate sensor that can be read non-destructively, then none of this modulus stuff is necessary anyway. Read once to set the proper gain on the ADC, then read again to get the fine bits. That way you still get the "unbounded HDR", but you don't need an ADC for every single receptor site, and you don't waste space on the counter or comparator. Instead, all of that lives on the image processor where it's not going to cause heat noise.
That's such an incredibly enormous hand-wave to me that I feel I've got to be missing something. Why make it so complex with all the modulus stuff if we have a magic sensor?
And we haven't even gotten into issues like thermal noise... Who knows, it may be great, but right now it just looks like a press release :-)
Oh the 1/10,000,000th comes from assuming they need to reset 128 * 8000 times per second to merely equal current sensor performance before losing sensitivity to misc losses.
If they want one more bit of DR they need to cycle in half as much time, and so on. Similarly, if we generously assume they're only losing 50% efficiency to cycle time and real estate then double again.
Also don't forget that you actually need to cycle significantly faster than just 128 * 8000 times per second, since presumably you can't integrate light into the capacitor while you're also zeroing it. Whatever goes into the capacitor while it's draining is lost, because it drains to zero.
Seems like a better approach would be to add a "gain pixel" to traditional sensors.
So imagine you've got your Bayer grid (or a layout that functions similarly), one of the pixels in the grid gets read out first. That one sets the gain for the ADC for the surrounding pixels. The downside is your iteration couldn't be linear and the gain control would be a lot more complex, but you could do it with technology that actually exists, and you wouldn't lose half of your vertical resolution by doing it a whole line at a time like Magic Lantern's dual-ISO mode.
You'd have to do fancy processing around high-contrast edges, but what else is new.
All good points. (And my figure actually assumes 10x faster and no loss of time between clearing wells and capturing photons -- i.e. 90% of time is spent capturing photons and 10% is spent clearing the well with no downtime -- to be generous. And, again, just to equal current performance.)
The underlying idea is sound -- storing bits is a linear problem but storing photons is an exponential problem. Of course you're still going to have to dump exponential amounts of energy... (So forget the "unbounded" part.)
Yeah I was thinking that too. No matter how tiny a capacitor, you have (tens of) millions of them on the sensor, one per photoreceptor, and you're charging and discharging them 100k-10M times per second. That's a lot of energy being accumulated and dissipated right next to the photoreceptors. Thermal noise is gonna be nuts on a modulus sensor, too.
No system is ever unbounded, period, unless you live in a world of frictionless spherical cows. At some point you run into issues caused by accuracy, energy, capacity, etc.
Consider that writing to a DRAM cell -- which involves charging or discharging a capacitor -- takes a few nanoseconds. This is at least four orders of magnitude faster than your estimate of 100µs.
Actually they're not so much faster -- although the camera companies don't have the economies of scale to iterate nearly as fast as the phone companies (so they hang onto a processor generation for a couple of years).
A good smartphone can capture 4K video or 20fps 8MP or whatever off the sensor — both around 160MB/s. A $10k DSLR is handling 16MP at 11fps — that's around 300MB/s (remember that the DSLR is handling 14-bits per sensel).
The interesting thing is that if you look across the camera lineups, there's almost no difference in CPU between a $300 consumer camera and a $10k professional (there is a big difference in RAM - the DSLR is able to store 50+ uncompressed images in RAM (let's say 3-4GB), while the phone or low end camera processes them into JPEG before storage.
You might be conflating cpu processor speed with, what is essentially "io" (light as an input, I'd suppose) and basically waiting for physics to happen.
Comments
The actual headline of the article is actually more informative and less misleading than the subhead:
"Unbounded High Dynamic Range Photography Using a Modulo Camera"
In photographer's terms, they're referring to ending blown highlights (caused by filling the wells in a sensel) not saturated images (caused by exaggerating color data).
There's a school of photography that basically lives and dies by over-saturating images (e.g. http://kenrockwell.com ) and they'd likely have a cow over the headline, but understand and be totally onboard for the actual goal.
Depending on how quickly the well can reset and how many times resets can be counted this will either be better or worse than simply improving existing approaches. The state of the art in full frame sensors is just under 15-bits of dynamic range. So to equal it an 8-bit sensor would need to be able to count resets at least 128 times AND be able to do this in 1/8000 of a second, at the same sensel density and efficiency! Don't hold your breath.
http://www.dxomark.com/Cameras/Nikon/D810
(I'd also suggest that trends towards using multiple sensors to assemble high resolution images will easily blow this away since they can use high- and low- sensitivity sensors to synthesize more dynamic range and resolution. Also see recent Olympus patents to capture polarization data at the same time.)
BTW: assuming they can do this stuff, they presumably can read the sensor pretty darn fast — so they should be able to eliminate image "tearing" in digital video and the need for mechanical shutters altogether. That's probably a bigger issue than dynamic range.
I suspect that clearing a well in < 1/10,000,000 of a second (which would only lose 10% of the photon detection time during a 1/8000s exposure) is going to be tough. Assuming 25% of the sensel real estate needs to be sacrificed to the modulo circuitry, that's a loss of 32.5% of photon data which is about the same loss as for pellicle mirrors (as seen in Sony's pretty unsuccessful SLT cameras) for a feature that won't be much use in many situations.
Actually, I suspect that even that crowd, after the initial shock would be absolutely on board. This would mean that no longer would the sensor control how the high-end clipping happens, but now that could be tailored specifically via a post-processing filter. E.g. to exactly match a long-gone film emulsion's behavior, "enhance" a classic, or to simply to create new clipping functions.
This is a bit analogous to the flexibility that black and white shooters gain by working with modern color sensors. When shooting B&W film, the tonal result was controlled via a combination of the film's spectral response and any colored lens filters applied, the latter used to manipulate tonal contrast between different colored subject matter (e.g. sky, plants, skin tones, etc.). Now a skilled photographer can shoot color RAW w/ B&W preview (for a preview of the image luminosity) then adjust in post to nail the image effect without having to mess even thinking about what colored lens filter(s) to have used.
Good point -- if this worked even moderately well -- e.g. it lets you capture 20 bits of DR -- a photographer could go crazy with the old film approach of "expose to the right" and simply keep settings such that every shot is "over" exposed and simply fix it all in post (it's pretty hard to over-expose by more than 6 stops unless you're trying...)
To emulate film stock we need a solid 14-bits of Dynamic range (while DxO says the D810 has it, the article below says 11-bits, but the best video cameras are getting there):
http://wolfcrow.com/blog/where-cameras-stand-in-dynamic-rang...
Ideally we'd get something like this AND capture polarization data as well (per the recent Olympus patent). Then you can apply polarization and exposure correction in post.
I identified the same issue with cycle time, but to me the biggest problem is the assumption of an accurate, sensitive, non-destructive-read sensor. They're assuming that they can charge a capacitor from a photodiode accurately, and also hook a voltage comparator (which is effectively the same as an ADC) to it constantly to get an accurate pulse to count.
Assuming you have an accurate sensor that can be read non-destructively, then none of this modulus stuff is necessary anyway. Read once to set the proper gain on the ADC, then read again to get the fine bits. That way you still get the "unbounded HDR", but you don't need an ADC for every single receptor site, and you don't waste space on the counter or comparator. Instead, all of that lives on the image processor where it's not going to cause heat noise.
That's such an incredibly enormous hand-wave to me that I feel I've got to be missing something. Why make it so complex with all the modulus stuff if we have a magic sensor?
And we haven't even gotten into issues like thermal noise... Who knows, it may be great, but right now it just looks like a press release :-)
Oh the 1/10,000,000th comes from assuming they need to reset 128 * 8000 times per second to merely equal current sensor performance before losing sensitivity to misc losses.
If they want one more bit of DR they need to cycle in half as much time, and so on. Similarly, if we generously assume they're only losing 50% efficiency to cycle time and real estate then double again.
Also don't forget that you actually need to cycle significantly faster than just 128 * 8000 times per second, since presumably you can't integrate light into the capacitor while you're also zeroing it. Whatever goes into the capacitor while it's draining is lost, because it drains to zero.
Seems like a better approach would be to add a "gain pixel" to traditional sensors.
So imagine you've got your Bayer grid (or a layout that functions similarly), one of the pixels in the grid gets read out first. That one sets the gain for the ADC for the surrounding pixels. The downside is your iteration couldn't be linear and the gain control would be a lot more complex, but you could do it with technology that actually exists, and you wouldn't lose half of your vertical resolution by doing it a whole line at a time like Magic Lantern's dual-ISO mode.
You'd have to do fancy processing around high-contrast edges, but what else is new.
All good points. (And my figure actually assumes 10x faster and no loss of time between clearing wells and capturing photons -- i.e. 90% of time is spent capturing photons and 10% is spent clearing the well with no downtime -- to be generous. And, again, just to equal current performance.)
The underlying idea is sound -- storing bits is a linear problem but storing photons is an exponential problem. Of course you're still going to have to dump exponential amounts of energy... (So forget the "unbounded" part.)
Yeah I was thinking that too. No matter how tiny a capacitor, you have (tens of) millions of them on the sensor, one per photoreceptor, and you're charging and discharging them 100k-10M times per second. That's a lot of energy being accumulated and dissipated right next to the photoreceptors. Thermal noise is gonna be nuts on a modulus sensor, too.
No system is ever unbounded, period, unless you live in a world of frictionless spherical cows. At some point you run into issues caused by accuracy, energy, capacity, etc.
maybe that energy can be used to recharge the battery
They don't count resets. They recover the number of resets via clever post processing.
Consider that writing to a DRAM cell -- which involves charging or discharging a capacitor -- takes a few nanoseconds. This is at least four orders of magnitude faster than your estimate of 100µs.
I don't think it's a problem.
What's funny is that the iPhone (and a high end Android phone) has way faster processors than your typical $10.000 DSLR.
Actually they're not so much faster -- although the camera companies don't have the economies of scale to iterate nearly as fast as the phone companies (so they hang onto a processor generation for a couple of years).
A good smartphone can capture 4K video or 20fps 8MP or whatever off the sensor — both around 160MB/s. A $10k DSLR is handling 16MP at 11fps — that's around 300MB/s (remember that the DSLR is handling 14-bits per sensel).
The interesting thing is that if you look across the camera lineups, there's almost no difference in CPU between a $300 consumer camera and a $10k professional (there is a big difference in RAM - the DSLR is able to store 50+ uncompressed images in RAM (let's say 3-4GB), while the phone or low end camera processes them into JPEG before storage.
You might be conflating cpu processor speed with, what is essentially "io" (light as an input, I'd suppose) and basically waiting for physics to happen.