Color Pseudoscience

Day after day, year after year, I am flummoxed by remarks like these about cameras:

  • It needs better color science
  • This one doesn't render skin tones as well as the other
  • I wish it had better highlight roll-off
  • Bring back the Panny mojo!

Let's assume your camera records RAW, compressed RAW, or at least 10-bit 4:2:2. If your camera records only 8-bit 4:2:0, I can understand (sorta).

Let's review how an image sensor works. Light strikes a matrix of metal, exciting a photodiode to convert those photons into electrons. At this stage, the graph of photon-input to electron-output is linear. If 2,000 photons would make 1,000 electrons, then 10,000 photons would make 5,000 electrons. (Somewhere along the way, this count of electrons [charge] is translated into electron pressure [voltage], but that doesn't matter here.)

The level of electrons is counted, on a scale of 1 to 4,096 (if 12-bit). We say at this point that the image was digitized (because we're using digits, get it?). We can then save this number to disk. Chances are we are too stingy with our disk to save it straight away. So we divide the number by 4 (10-bit) or even 16 (8-bit), and then save that smaller number to disk. No, that's not enough, we might collect a bunch a numbers together and replace them all with a simpler pattern of numbers. This is called compression.

At this point I'm not sure where the mojo is coming from, because we're just scientifically measuring the light at different points in the frame.

All that mojo comes later. It's there, I do not dispute that one camera's default output looks different than another's. That's because processing is inserted between the reading off the sensor and the writing onto disk. I referred to all that as "processing," which as most of you know encompasses different things: changing color, changing contrast, noise reduction, etc.

But here's the thing. Any of that magical processing done in camera, is also possible to duplicate, or even counteract, on your computer.
 
But here's the thing. Any of that magical processing done in camera, is also possible to duplicate, or even counteract, on your computer.

But then there are ARRI ALEXA style sensors that are using dual gain stages in the sensor to create a 16-bit HDR image in real-time with extremely wide dynamic range. I don't think there's any easy way to recreate this in post.
 
Short answer to last line: nope (until there's advanced AI to fix all the problems in a realistically timely manner). Or are you doing an Eric-style troll/joke?

Most cameras suffer from exposure variant color- a major pain to fix in post (shadows, mid-tones, highlights all having different color processing- it's a sloppy mess). The camera which does the best in the general case is ARRI (that's why it's still #1 after all these years. People aren't paying those prices for any reason other than those cameras provide the best color and DR: less work in post for the general case (and doing the best in mixed lighting).

Time is money.

You might want to start a conversation on the ARRI forum WRT color science as pseudoscience with ARRI color scientists (if you're being serious and this post is not a joke):
https://www.arri.com/en/learn-help/technology/image-processing
Film-like, Organic Look
Our camera’s unique color processing was developed by the same ARRI color scientists that worked on the ARRILASER and ARRISCAN. Therefore, they are intimately familiar with both film and digital color science. The ALEV sensor renders natural colors, great looking skin tones, and shows accurate color separation (important for green screen and other VFX work) while also demonstrating the ability to resolve mixed color temperature sources.
https://forum.arri.com/
 
Last edited:
Well - I have two Pentax DSLR, used Pentax for years - and one has very nice colour rendition from stage lighting and the other sucks! I cannot make one look like the other no matter how I tweak on Photoshop. I can get close on most colours, but the pinks and magentas are badly mangled and I cannot pull out a fuscia colour from magenta on one, but on the other it sees this without me needing to interact.

I think it's like printers - some, with the same colour inks just struggle and others shine.
 
Just to be clear, this is not a joke.

Take the Arri for example. Everything it is doing can be done in post, especially if you have a raw image, if the camera has the same dynamic range.

Arri has a dual-gain amplifier. An amplifier does not increase dynamic range. It boosts noise as well as signal. All that Arri does is a convenience.

Also, Arri is not the only camera with a dual-gain amplifier. For example, the original Blackmagic Pocket's sensor had one.

I don't understand the continued devotion to the Arri, a camera that came about 10 years ago. Yes, it has superb hand-crafted German engineering, but the sensor was not made in Germany. They bought it.

It still looks great, but cameras since then have matched or equaled its dynamic range. And so I see no reason why it would be superior.

---

I guess I'm not saying it would be easy to grade two cameras to look exactly alike if you lined up their frames and switched back and forth and zoomed in 100x. But the difference would be so subtle as to be unnoticeable at normal viewing.

The difference between the Arri and say, the FS7, is less than the difference between the FS7 by default and the FS7 with a LUT. Also there are several videos comparing, say the Alexa to the Pocket, and grading them so close that you can't easily identify each one.
 
I've worked with footage from all top cameras down to budget consumer cameras. I've also written sophisticated 3D color processing software (CPU & GPU) and spent way too much time analyzing and fixing skin tones (e.g. the Filmic Skin project for Canon DSLRs and Cinema EOS cameras). Frankly, I don't really want to deal with obvious defects in color science anymore. Currently ARRI still has a comfortable lead in color science / color and image processing. Doesn't matter what country sources what part or wrote the software: ARRI deals with color better than anyone else, and their over/under exposure recovery is still better than cameras that claim the same or higher DR. This is what you can do with exposure invariant color science- some cameras might provide decent e.g. over exposure recovery in terms of clipping, however the colors are messed up (especially skin tones). Same thing for shadow recovery. Professionals are more bottom line regarding costs than hobbyists- if Red or Sony or Canon or Panasonic was in the same ballpark regarding DR and color, no one would waste $100+K on Alexa packages.

We've seen studio tests where all the cameras are matched to the Alexa. Better tests would be to let each team try to create accurate color as the viewed by the eye, then compare results. In this case cameras like the FS7 will tend to look significantly worse than the Alexas, especially in mixed lighting (this is from viewing FS7 footage vs. ARRI in the wild).

Color correcting a single raw frame (as with still photography), where the exposure variant color isn't a huge deal isn't the same as correcting raw video, especially with mixed lighting as the camera moves around: this is a big deal and a major time burner to fix.

ARRI is simply more color accurate vs. the lower end cameras. For those with limited budgets, they'll be spending more time in post, and if the people doing post are paid hourly, the producers would be well informed to do some math vs. buying/renting an Alexa and saving $$$ in post fixing messed up / inaccurate color.

On a similar note: modern cameras should have a built in high quality colorimeter (e.g. + a separate color sensor so lens color corrections can be applied too if desired). Calibrated color straight from the camera will greatly improve final image quality- this has been possible for many years: guess they're saving this for a future product feature.
 
Take the Arri for example. Everything it is doing can be done in post, especially if you have a raw image, if the camera has the same dynamic range.

Arri has a dual-gain amplifier. An amplifier does not increase dynamic range. It boosts noise as well as signal. All that Arri does is a convenience.

Also, Arri is not the only camera with a dual-gain amplifier. For example, the original Blackmagic Pocket's sensor had one.

I haven't seen any other camera that is able to match either the color or DR of the ARRI ALEXA cameras. Sony and Panasonic can get close, but ARRI always has more DR than anyone else.
 
  • Like
Reactions: jcs
How blue is blue in comparison to pink? How saturated is red in the highlights vs the shadows?

These are all things that cameras do, but are painstaking to fix in post.


So, colour science or psuedo colour science, or even popular opinion, the cameras matter. Look at how good the Varicam is, and yet it is not typically in the colour science conversations. Or the fact that “Arri” is only a buzz word the last 4years, when the camera is older than that.

Opinion and perception are weird factors. Someone did a test for mojo in stills cameras, and in a loaded blind test, people favoured the Sony.

The term “colour science” may be a misnomer. Kind of like talking about DR. But the thing it represents is generally understood, despite being vague and not exact.
 
Random thoughts.

In 2003 I had a 16bit 6k raw hasselblad CCD still camera.. yes it could make extrodinary images.. but the software was so awful that what was possible on a commercial deadline made it very very very hard to use. Workflow does matter.

Arri win on the basics.. big pixels and raw.. clearly it is going to be a winner. As we see the only way to make it better was to double the chip size (still smaller than that hasselblad)

Let's assume your camera records RAW, compressed RAW, or at least 10-bit 4:2:2.

Its a duff assumption many cameras that these allegations (It needs better color science This one doesn't render skin tones as well as the other I wish it had better highlight roll-off)

Dont have those facilities.

But here's the thing. Any of that magical processing done in camera,.. (can be) counteract(ed)

Not so - once you throw two colours into one 'bucket' that 'choice' is burned and cannot be undone - a computer cannot uncompress compressed footage

At this point I'm not sure where the mojo is coming from, because we're just scientifically measuring the light at different points in the frame.


Sensors deliver a huge amount of data, how it is 'written' usually includes lots of compression - mapped - squashed. This process can be done with or without mojo.

--

I remember seeing a sony shot of football (soccar) - it was horrific - but the camera had been 'painted' to cover the bright sun on one half of the stadium and deep in the shadows on the other half.
Noone would use that profile for 'cinema' but it really did the job - so in effect it was horrific to some.. but no football fan wants not to see/uderstand what is occuring.. it had mojo to a footbal fan but no mojo for a cinematographer with some control of the input light/contrast ratio.

Overall I would agree that most raw/12bit files should be lutted to basically match - as Yeldin suggests
 
Color science is a real science, it's super complicated mathematically (unfortunately) and it also includes human perception (which is even more complicated (fortunately ;)). Viewing calibrated color on a high quality calibrated monitor is only a small piece of the puzzle. A good colorist knows to check the scopes, skin tone line etc., and to periodically look away from the screen to avoid 'color retention' in the eye-brain perception system. Meaning, you can stare at really bad color, and after a while it can look OK because the brain does some fancy fix-up filtering. Take a break and come back later and it can look terrible (after the eye-brain system has been reset).

From camera to computer to NLE/post for calibrated color work to viewing on a phone, TV, or computer monitor is an incredibly complicate chain of color science. A break anywhere in that chain results in messed up color. Many here I'm sure have been annoyed by 0-255 / 16-235 and incorrect matrix bugs in e.g. QuickTime, ffmpeg, YouTube etc.

Another thing experienced colorist do is view the content on all the target devices (phone, TV, monitor, local devices, online services) as sometimes one or more device will need an adjustment that when done correctly, will look pretty good on all devices and viewing targets. It's a lot of work and starting with a camera with good (or great!) color science significantly reduces the amount of work. Which brings up the concept of color latitude in post: ARRI does the best here too- when a colorist does the multi-device test, ARRI sourced material is going to look better, because it's more accurate which ultimately leads to less non-linear color errors which pop up with lower quality / less accurate color processing cameras. NLE/color-tool bugs can also mess up the color pipeline.

+ I think it's Steve Yedlin, ASC.
 
I'm really scratching my head, because it still doesn't make sense. I implore you not to write me off as too stupid. I got into filmmaking 30 years ago, I made A's in school, I did video production professionally for several years, got burned out, started over and taught myself programming, which I am now paid to do. But I can't understand this.

Let's go back to where we agree. In fact, let's shelve Arri for the moment and just think of something calmer, like Sony vs. Panasonic.

  • If you're talking about quick and easy workflow, and you just want a camera that gives you the look you want out of the box, I get it.
  • If you're saying that sometimes you can't make two cameras look exactly the same no matter what you do, I will agree.
  • There are cameras where I don't like the default look. Almost all cameras, actually.
  • There are videos on Vimeo and Youtube where I don't like the highlight roll-off or color. Almost all videos, actually.

But:

  • I was under the impression that almost everyone was running their footage through DaVinci Resolve (or whatever). Almost no one is just handing off their raw footage or making cuts-only editing anymore. Even if you don't have to, it's just too tempting with all those knobs.
  • Even if you are in a hurry, I thought you could establish presets after getting it right once for your camera, make a LUT, whatever.
  • These differences that you are talking to me about seem like very small differences, not visible to the unaided eye. If you're just being a perfectionist, just trying to please yourself, fine. (I want global shutter. Who can see the difference between rolling and global shutter in most situations?) But if you're saying that customers will notice the difference, if you're saying that if you buy Camera B, it will affect the enjoyment and business you could have got with Camera A, and you can never make Camera B look good enough for the people who like Camera A (average viewers, again, not pixel peepers), then I am lost. I just don't see it when I look at demo videos of all the professional cameras and DSLRs that have come out in the past 3-5 years.
 
Last edited:
Day after day, year after year, I am flummoxed by remarks like these about cameras:

  • It needs better color science
  • This one doesn't render skin tones as well as the other
  • I wish it had better highlight roll-off
  • Bring back the Panny mojo!

Let's assume your camera records RAW, compressed RAW, or at least 10-bit 4:2:2. If your camera records only 8-bit 4:2:0, I can understand (sorta).

Let's review how an image sensor works. Light strikes a matrix of metal, exciting a photodiode to convert those photons into electrons. At this stage, the graph of photon-input to electron-output is linear. If 2,000 photons would make 1,000 electrons, then 10,000 photons would make 5,000 electrons. (Somewhere along the way, this count of electrons [charge] is translated into electron pressure [voltage], but that doesn't matter here.)

The level of electrons is counted, on a scale of 1 to 4,096 (if 12-bit). We say at this point that the image was digitized (because we're using digits, get it?). We can then save this number to disk. Chances are we are too stingy with our disk to save it straight away. So we divide the number by 4 (10-bit) or even 16 (8-bit), and then save that smaller number to disk. No, that's not enough, we might collect a bunch a numbers together and replace them all with a simpler pattern of numbers. This is called compression.

At this point I'm not sure where the mojo is coming from, because we're just scientifically measuring the light at different points in the frame.

All that mojo comes later. It's there, I do not dispute that one camera's default output looks different than another's. That's because processing is inserted between the reading off the sensor and the writing onto disk. I referred to all that as "processing," which as most of you know encompasses different things: changing color, changing contrast, noise reduction, etc.

But here's the thing. Any of that magical processing done in camera, is also possible to duplicate, or even counteract, on your computer.


Let's not forget about the very front of the chain(at least IN the camera): The CFA in front of the sensor.
 
When you have a deeper understanding of what color is (in volumetric 3D space, and how linear & non-linear transformations and rotations affect color (for better or worse)), and how the eye-brain system processes color, with especially high discrimination for skin tones (evolutionary optimization (or for those that don't believe in that: an optimization/adaptation for health/reproductive cues)), the question of whether color science is real or not becomes clear. Lenses have MTF (+ they act as a color filter too), CFA's (as R&G noted) affect color, sensor design affects color, processing in camera at many layers, compression affects color too- all the way to the final moment when the content is played back and perceived by another consciousness (human, dog, cat, AI whatever ;)), that's color science.

A simple visualization experiment: when ARRI captures a frame, we can say that the resulting color space (3D cube) has well separated colors (little or no cross-talk meaning R, G, B are well separated, with no colors polluting each other (see ARRI's comment about that in prior link- it's super important!), and little or no volumetric distortions or rotations. In my experience Sony & Panasonic (and a to a lesser extent, Canon), can provide an initial image which can look OK, but then after processing or grading, ugly issues become visible, especially with skin tones (because the eye-brain is very fine-tuned there). This is because the captured volumetric 3D colorspace of the image is distorted in such a way that additional processing amplifies the initial distortions which must then be corrected with non-linear and sometimes localized post tools (masking/windows, selective color repair, etc.). When I say distorted, that is distorted from what would otherwise be a correct 3D color volume.

To simplify viewing, 2D mapping is sometimes uses, e.g. the CIE color space (from: https://medium.com/hipster-color-science/a-beginners-guide-to-colorimetry-401f1830b65a ):

http://tny.im/keY
keY


Or a 3D to 2D projection of a spectral locus:

http://tny.im/keZ
keZ


There is a TON of information on color science available online- just gotta spend some time doing research. CIE, SMTPE, ACES (IDTs), Rec601/709/2020, JPEG & MPEG (groups), again, very complicated and lots and lots of info. All of it together is color science (a combination of math, physics, and human eye-brain perception).

Have you ever tried calibrating a monitor? Observed the before/after- it's pretty dramatic!

How about a printer!! (or tried to get accurate skin tones printed with a new printer for the first time (at home or a Kinkos/pro-printing company etc.)?!?

This is why I keep mentioning that cameras really need accurate color science at the camera level, including new calibrated color sensors.

Also try the hand-mirror + calibrated monitor live test with your camera: attach camera to accurate/calibrated monitor, observe your own face (or a model's) then compare to what you see with your eyes. This can be quite an eye-opening experience! Now record material, and playback on the same device and observe self in hand mirror or model. In my experience we have a long way to go to get accurate color as the eye sees it from lens+camera to monitor to final perception in the mind.

A really well calibrated/tuned image, when displayed on a variety of mis-calibrated devices, will still look good because it's at the 'center' of the 3D volume in a sense, so device distortion won't really f' up the final image. This exact same pattern applies to audio/sound, and why it's important to listen on neutral monitors as well as crappy iPhone/Android headphones and really great headphones (e.g. electrostatics and anything you might have on hand in between).
 
Last edited:
Let's review how an image sensor works. Light strikes a matrix of metal, exciting a photodiode to convert those photons into electrons. At this stage, the graph of photon-input to electron-output is linear. If 2,000 photons would make 1,000 electrons, then 10,000 photons would make 5,000 electrons. (Somewhere along the way, this count of electrons [charge] is translated into electron pressure [voltage], but that doesn't matter here.)

The level of electrons is counted, on a scale of 1 to 4,096 (if 12-bit). We say at this point that the image was digitized (because we're using digits, get it?). We can then save this number to disk. Chances are we are too stingy with our disk to save it straight away. So we divide the number by 4 (10-bit) or even 16 (8-bit), and then save that smaller number to disk. No, that's not enough, we might collect a bunch a numbers together and replace them all with a simpler pattern of numbers. This is called compression.

At this point I'm not sure where the mojo is coming from, because we're just scientifically measuring the light at different points in the frame.


You've left off a few key parts.

The CFA, the de-mosiac algorithm, the color Matrix table, the clock speed (integration time), the calibration techniques and choices. The choice of IR filter and choice of OLPF could also arguably affect spectral response. These choices are all PRE-RAW or baked into the RAW.

There's a problem with the use of RAW in the lexicon, It infers that it's RAW sensor data when it's really not.

There's a CFA. That colour filter array is tuned in some way to reproduce colour. The colour of those filters is something that's highly variable and it's still a CHOICE being made by someone that's not you and not something you can normalise or override easily.

That colour is also "created" through the use of a demosaic algorithm, all the maths that interpolates colour from the photosites. Again, this is variable between manufacturers. It haven't done the tests for a while, but or a long time it was accepted practice to use a RED ROCKET to speed up grading in Resolve, but or the final render you'd turn off the RED ROCKET card for final render / output on big movies and and use the internal Resolve de-bayer because it produced a better result when you didn't need the speed of playback. That's because there's two implementations of a de-byaer. You can also test this on Sony RAW footage in resolve as well where you have a choice in the type of algorithm to use. Many advanced stills imaging processes allow even more customisations of the debayer algorithm choice.

When you get a lot of cameras that seem to have the same sensor origin, one of the reasons they make different pictures is because each manufacturer using that base sensor is then adding their own CFA.

The matrix is basically a recording or table of subjective color choice. It's interpretive. It's not scientific. I think of it as a manufacturer's LUT that is baked into the RAW. You can't really change what that is either.

The speed with which you operate a camera and sensor (like a CPU hz) also affects how fast the sensor clocks out frames, also known as integration time. This directly affects motion cadence and motion perception.

The IR filter stack and use of OLPF also affects the way the colour is reproduced. If a weak IR filter is used then it will produce a very different result to a camera that uses a stronger IR filter, especially around skin tones.

There's also a lot of calibration that happens because it's not unusual to have a lot of sensor to sensor variation, and when your'e trying to create a consistent look from a camera, there needs to be a way to adjust and calibrate every single sensor towards that norm, an internal standard of how a camera reproduces a spectral response.

All of these steps are intrinsic to a manufacturer and make colour become a "science". When I speak to the engineers that design cameras, one of lead's uses "air quotes" when he talks about colour science because it is kind of a consumer / end user term that means nothing to him, but it actually IS the sum of these many other choices that adds up to subjective judgements about the way a camera "represents" colour and motion.

I think it's a simplistic logic to assume RAW means sensor data in a simple mathematical table that is somehow pure when it gets to to the end user in a RAW processing application. There are many choices made by someone else that alter what numbers end up in that RAW container before it gets to you. Some of those you can correct and adjust for, some are always going to be intrinsic.

JB
 
Last edited:
Why make it more complicated than whats needed? Know your tool, expose correct and select the right white balance.

Agreed!

Shot on F55, Golman chose it over the ARRI and the Venice for season 3. I wouldn't enjoy it more or less if it was shot on an ARRI or a piece of sting. Looks good and it's nicely shot. Production values make it what it is.

https://www.studiodaily.com/2018/06...man-asc-abc-capturing-lavish-details-royalty/

Your average viewer wouldn't know what an ARRI was. Often we are our own worst enemy. Odd to see it reconstructed but I was filming at the presser shown at 44 seconds, 21 October 1966. Was one of the first cameraman there.

https://www.youtube.com/watch?v=vLXYfgpqb8A

Chris Young
 
Last edited:
I was under the impression that almost everyone was running their footage through DaVinci Resolve (or whatever). Almost no one is just handing off their raw footage or making cuts-only editing anymore. Even if you don't have to, it's just too tempting with all those knobs.

You need to get out in the real world more. We hand off a considerable amount of footage that does get cuts only with perhaps a title card and then is broadcast or more often, posted to YouTube or a corporate web page.
Also, EPK, Press Junkets, the footage is almost always immediately used and mostly without any CC. What about Live Streaming too? I was one of five camera ops for the Facebook Live stream for the NAACP Awards in the
Spring. What came out of the cameras (no engineer or painting going on) went straight out over Facebook. In all of these cases, what you see is what the audience gets.
 
Last edited:
In the case of handing over the cards, I would think that the studio would adjust it to their taste anyway.
In the case of a studio too hurried to do so, I would think that they could apply a LUT with a click.
In the case of streaming, I would think you could apply a LUT to the stream before it leaves your computer.
In any case, I would think you could upload a LUT into your camera. Is that possible yet? I thought some cameras let you do that now.
 
Last edited:
In the case of handing over the cards, I would think that the studio would adjust it to their taste anyway.
In the case of a studio too hurried to do so, I would think that they could apply a LUT with a click.
In the case of streaming, I would think you could apply a LUT to the stream before it leaves your computer.
In any case, I would think you could upload a LUT into your camera. Is that possible yet? I thought some cameras let you do that now.

You have all of these replies that are "could" and "should". I'm just saying, in the real world, a lot of the anal retentiveness we have about these details goes out the window.
Most of these type clients (EPK, Studio Publicity, Home Entertainment) don't want to deal with log, LUTs or CC. That's just a fact. I shot a feature documentary that was shown at Cannes this year in REC 709. They did do some CC at least
but I knew going in that they didn't want to shoot Log, RAW or anything unusual because when we did camera tests, they told me. Shot it Canon WDR and they were happy with the footage. It's their footage so I just do
what they say they want. Hallmark Channel this year was the same. Straight out of camera, onto the web.

It's similar to when we hire sound mixers and they have these killer sounding files recorded on a $6k Sound Devices or Zaxcom recorder, yet the relatively crappy camera audio recorded with a usually lower end or outdated wireless system is almost ALWAYS what actually gets used.
The C200 as far as audio, is actually one of the better sounding cameras, but it still pales in comparison to what a 633 or an 833 records. Drives me crazy but a lot of our clients are just too harried and on deadlines to do audio conforming to use the good stuff. It's just for backups in
case the crappy receiver on the camera has a dropout or the receiver batteries die and I fail to notice it for a few minutes as I am shooting. This is real world production.
 
  • Like
Reactions: jcs
Back
Top