Acutance

I. Definition

Acutance is an objective physical quantity used to measure image clarity. By contrast, sharpness is the subjective human perception of image clarity; the two are distinct concepts. Early acutance calculations were based on the gradient of an edge profile: the steeper the transition, the higher the value. However, an edge gradient describes only the physical rate of transition and does not account for the fact that images are ultimately perceived by the human visual system. Human contrast sensitivity varies nonlinearly with spatial frequency, and the perceived acutance of the same image also depends on viewing conditions such as display size and viewing distance. These factors were not included in earlier methods.

For this reason, modern standards such as ISO 12233:2023 Annex L and IEEE 1858 CPIQ redefine acutance as a single value obtained by weighting and integrating the system spatial frequency response (SFR) with the human contrast sensitivity function (CSF) under specified viewing conditions, making it more closely aligned with perceived sharpness. SFR can be obtained using a variety of test patterns. ISO 12233 originally defined only the slanted-edge method, but later progressively incorporated alternatives such as the Siemens star and dead leaves chart.

II. Calculation Principle
According to ISO 12233:2023 Annex L, acutance uses the human contrast sensitivity function (CSF) as a weighting function to integrate the system spatial frequency response (SFR) data, ultimately yielding a quantitative value (Q value) associated with subjective sharpness perception.

Using the CSF as the weighting function, the SFR data are weighted and integrated to obtain the Q value:

$$Q = \frac{\sum_{i=1}^{N} SFR_i \, CSF_i}{\sum_{i=1}^{N} CSF_i}, \quad i = 1, 2, \dots, N \tag{L.2}$$ where \(SFR_i\) is the Spatial Frequency Response value at the \(i\)-th spatial frequency; \(CSF_i\) is the human contrast sensitivity weight at the \(i\)-th spatial frequency; and \(N\) is the number of spatial frequency points included in the calculation.

The three key elements involved in Q-value calculation are described below.

1.Contrast Sensitivity Function (CSF)
The CSF model originates from the S-CIELAB perceptual color space and has been adopted by the ISO 12233 standard for perceptual weighting of the luminance channel.

The CSF curve (see figure below) describes the human visual system’s sensitivity to contrast at different spatial frequencies:

  • Horizontal Axis: Spatial frequency $f$, in cycles per degree, representing the fineness of image detail.
  • Vertical Axis: contrast sensitivity; higher values indicate that contrast differences at that frequency are more easily perceived by the human eye.

This curve shows that human vision is most sensitive to low-to-mid frequency details at around 4 cycles per degree, while sensitivity decreases significantly for both very high and very low spatial frequencies.

Its mathematical model is:

$$csf_{\text{lum}}(f) = \frac{(a \cdot f^c) e^{-bf}}{K} \tag{L.1}$$

The values of the parameters are explicitly specified in ISO 12233:2023 Annex L: \(a = 75\), the amplitude coefficient; \(b = 0.2\), the high-frequency attenuation coefficient; \(c = 0.8\), the frequency exponent coefficient; \(K = 102.16\), the normalization constant (which sets the peak value of the curve to 1.0); and \(f\), the spatial frequency in cycles per degree.

2. Viewing conditions

The calculated result of acutance is strongly dependent on viewing conditions: the same image may appear sharp on a small display but blurry when enlarged or viewed on a larger display. This difference is determined by the contrast sensitivity characteristics of the human visual system—under different viewing distances and display sizes, the same spatial frequency corresponds to different perceptual sensitivities.

To ensure that acutance values calculated by different devices and laboratories are comparable, ISO 12233:2023 Annex L explicitly specifies three sets of standard viewing conditions (VC1/VC2/VC3) (see table below) to unify key parameters such as display size, viewing distance, and pixel pitch.

3. Unit conversion

Before applying the CSF to SFR data, the frequency unit of the SFR must first be converted from cycles per pixel to cycles per degree to match the independent variable of the CSF. According to the Nyquist sampling theorem, the maximum representable frequency \(f_{cut}\) in an image is 0.5 cycles/pixel. This conversion requires the pixel size of the image as finally viewed by the observer (i.e., the pixel pitch \(p\)) and the viewing distance \(D\), and is performed using Equation (L.3):

\[ f_{\text{cycles/degree}} = \frac{\pi D}{180 p} f_{\text{cycles/pixel}} = \frac{\pi D N_H}{180 H} f_{\text{cycles/pixel}} \tag{L.3} \]

where \(N_H\) is the number of pixels in the vertical direction of the image; \(H\) is the image height; \(p\) is the pixel pitch of the display device; and \(D\) is the viewing distance.