Perceptual tolerance illustration

/

Insights

Perceptual Optimization in Video Coding: Approaches and Engineering Tradeoffs

Perceptual Optimization in Video Coding: Approaches and Engineering Tradeoffs

Where perceptual optimization fits in a video workflow—and how to balance compression gains, processing cost, control, and picture quality.

Where perceptual optimization fits in a video workflow—and how to balance compression gains, processing cost, control, and picture quality.

Latest update

Sergio Sanz

Posted originally

Sergio Sanz

At a glance

  • Viewers do not notice every detail equally: Perceptual optimization uses this insight to aim for lower bandwidth use while preserving picture quality, or better quality at the same bandwidth.

  • Start with the workflow, not the AI label: Optimization can act inside the encoder, before encoding, or on both the encoding and playback sides. Where it fits matters more than whether it uses AI.

  • Every option involves practical tradeoffs: Integration effort, processing cost, control, and picture quality all matter. Encoder-side optimization and preprocessing can leave playback unchanged; the “both sides” approach adds processing during playback.

  • More complex does not automatically mean better. Choose the approach that meets the workflow’s needs, and judge its benefits across the complete pipeline—not by how sophisticated it sounds.

Modern video codecs efficiently remove statistical redundancy, but not every part of a video contributes equally to what a viewer perceives. Perceptual optimization exploits this difference to seek lower bitrate at comparable perceived quality, or better quality at the same bitrate.

Given the significant attention AI receives today, it can be tempting to equate “AI-based” with “better” and overlook approaches that do not use it. For a broadcaster or video streaming service, however, the first question should not be whether a method uses AI, but where perceptual optimization can be introduced: inside an available encoder, before encoding without changing playback, or on both sides of the codec. Each choice has different implications for integration, computational requirements, and picture quality.

This article reviews representative methods at those locations and considers what their design choices mean for a production workflow.

What Is Perceptual Optimization?

Perceptual optimization uses knowledge of human visual perception to allocate, transform, reduce, or reshape visual information so that video can be compressed more efficiently while preserving perceived quality.

In practice, this often means protecting visually important features—such as text, faces, and strong edges—as well as uniform areas where banding and blocking artifacts can be particularly noticeable. In contrast, more simplification can often be tolerated in textured regions such as grass and foliage, or turbulent water—for example, splashes or waterfalls—where some changes are less noticeable. Film grain may also be reduced when it is not an intentional part of the look.

Where Can Optimization Fit?

The approaches reviewed here can be grouped by where perceptual optimization acts: encoder-integrated perceptual coding inside the encoder, external perceptual preprocessing before it, and neural codec wrappers around it, combining preprocessing before encoding with learned reconstruction after decoding. This placement has practical implications for integration with existing encoders and playback systems. The schematic below summarizes these deployment patterns.

Three-stage video processing infographic

Figure 1. Highlighted blocks show intervention points. Encoder-to-decoder arrows carry the bitstream. Deployment patterns can be combined. perceptual_optimization_approaches

Encoder-Integrated Perceptual Coding

Encoder-integrated methods bring perceptual reasoning directly into coding decisions [1].

  • Bit allocation and rate control spend more bits where distortions are easier to notice, for example by adapting quantization strength locally.

  • Rate-Distortion Optimization (RDO) selects coding options by balancing bitrate against distortion. Perceptually weighted costs can make that balance more relevant to viewers.

  • Transform and quantization decisions can preserve more visible components while coding less noticeable ones more coarsely.

Challenges and Tradeoffs

The main benefit of encoder-integrated perceptual coding is direct access to decisions that spend bits. With a fixed third-party encoder, however, customization is limited to the perceptual controls it exposes, whether through configuration options or a software development kit (SDK). These controls may not provide enough flexibility for a particular use case. Deeper changes require access to the encoder implementation. These encoder-side methods can preserve compatibility with standard decoders, but an optimization built into one encoder does not automatically transfer to another.

External Perceptual Preprocessing

External preprocessing modifies the input while keeping it directly displayable after conventional encoding and decoding. It can therefore be added without changing the codec or receiver [5]. This makes preprocessing easier to integrate into existing workflows, although its compression benefits can vary with the encoder, preset, and target bitrate.

Classical Deterministic Methods

Classical deterministic methods combine explicit signal-processing rules with models of visual perception. In the pixel domain, adaptive filters smooth more strongly where changes are less visible while protecting important structures. BilAWA/TBil filtering is one example [2], including a UHD HEVC extension [3].

Hand-designed filtering does not mean non-adaptive: processing can respond to local signal characteristics. However, designing and tuning these methods requires expertise in video signal processing and statistics.

In the transform domain, a mathematical transform such as the Discrete Cosine Transform (DCT) separates an image block into spatial-frequency components so that less visible components can be attenuated [4]. Just-noticeable-difference (JND) models estimate how much a signal can change before that change becomes noticeable.

A third deterministic branch is spatiotemporal filtering, which uses information across frames to reduce temporally uncorrelated noise or grain. It can improve compressibility, but motion-estimation errors, occlusions, scene cuts, and intentional grain require careful handling.

Challenges and Tradeoffs

The practical attraction of classical methods is the ability to inspect and control their behavior: an operator can impose explicit limits on filtering strength or pixel changes. For example, if preserving the source look is a priority, a mild deterministic prefilter can be designed with strict limits on the changes it introduces. Nevertheless, even small, bounded changes still need to be checked for visible artifacts. More sophisticated classical preprocessing methods may offer greater perceptual gains for a given bitrate, but can also carry substantial computational costs.

Fully Learned Preprocessing

Fully learned preprocessors use a neural network to predict a modified frame or a correction to it. Examples include Deep Perceptual Preprocessing (DPP) [5], IQNet [6], and HDR-JNDNet [7]. Their objective is to learn changes to the input that improve the tradeoff between bitrate and perceived quality.

The main attraction is content adaptivity: learned models can capture relationships that are difficult to express with hand-designed rules. More advanced networks could also use higher-level context—for example, recognizing turbulent water as a region where stronger simplification may be tolerated while protecting smoother areas such as the sky (see image below).

Challenges and Tradeoffs

The engineering challenge is translating this adaptivity into consistent improvements across diverse content. Training and tuning can require substantial computational resources, and the resulting model must preserve spatial fidelity and temporal consistency while avoiding excessive smoothing, unwanted detail, color shifts, and other visible artifacts. Strong results on individual frames are not enough if picture quality fluctuates during playback.

For real-time workflows, these quality requirements must also be met within the available processing and latency budgets. Where improving temporal consistency requires additional temporal processing, its computational and memory costs must be included in the deployment assessment. Training costs are one consideration; the recurring cost of processing each stream is another.

The practical question is whether the compression savings and quality improvements justify that operating cost.

Fountain image illustrating turbulent water and smooth sky

Figure 2. Turbulent water and smooth sky illustrate different perceptual challenges. This is a conceptual example, not a measured model result. hhi@water_vs_sky.png

Hybrid Learned-Control

This approach lets the neural network decide where and/or how strongly to process the video, while a deterministic filter performs the actual modification in a controlled and predictable way. Adaptive High-Frequency Preprocessing is one recent example of this pattern [8].

Challenges and Tradeoffs

Hybrid learned-control aims to combine the best of both worlds: the simplicity and control of deterministic filtering with the content adaptivity of learned models. However, the neural network can still select settings that remove important details. Without sufficient temporal context, it may select inconsistent filtering strengths for similar consecutive frames, making the image appear alternately softer and sharper. The network’s flexibility is also limited to what the underlying filter can do.

The practical question is whether the added adaptivity justifies the computational cost and complexity associated with the neural network compared with a well-tuned classical prefilter.

Neural Codec Wrappers

Neural codec wrappers place learned processing before encoding and learned reconstruction after decoding. Their attraction is that preprocessing and reconstruction can be optimized together to improve the tradeoff between bitrate and perceived quality [9]. The conventional codec can remain unchanged, but the intended end-to-end output depends on an additional receiver-side stage.

Challenges and Tradeoffs

A standard-compliant bitstream does not mean an unchanged playback workflow. This approach requires both the encoding and playback systems to run the corresponding learned models, meet their computational requirements, and support any necessary model updates. For existing receivers that cannot support the reconstruction stage, external perceptual preprocessing alone is easier to deploy.

Choosing for a Production Workflow

Start with what can change: If playback must remain unchanged, consider encoder-side optimization or external preprocessing. Neural wrappers become an option when both encoding and playback systems can support the additional processing.

Assess the complete workflow: For live video, evaluate whether the approach meets real-time requirements at an acceptable latency and operating cost. For offline encoding, processing time and cost per finished asset may matter more. Assess the complete pipeline, not just the optimization stage.

Protect the viewing experience: Test representative content, including difficult cases such as subtitles, grain, and scene transitions. Define acceptable changes, control processing strength, and provide a way to disable the processing when necessary.

Interpreting Reported Compression Gains

In this context, compression gains mean better perceived quality at the same bitrate, or a lower bitrate at the same perceived quality. A fair comparison evaluates the final decoded video with and without perceptual optimization under comparable conditions, combining quality measurements with viewing tests.

Published results are not always directly comparable because the content, encoders, settings, and quality measures differ. One study [10] reported better overall results for a learned prefilter than selected classical frequency-domain methods. However, it encoded frames independently, processed only brightness information, and included no viewer assessments. This does not establish a general ranking of classical and learned approaches. In the literature we reviewed, we did not identify broader like-for-like comparisons covering typical streaming and broadcast encoding, a wider range of classical methods, stability during playback, and viewer assessments.

Interestingly, the same study used classical preprocessing to help train its neural network [10], illustrating how classical and learned approaches can also complement each other.

Takeaway

Perceptual optimization is an engineering choice about where to intervene in the video pipeline and which changes are acceptable. The goal is to reduce information that costs bits but contributes little to perceived quality, without creating new visible problems.

The best approach is not necessarily the most complex one, but the one that delivers the strongest balance of compression and visual quality within the workflow’s processing, reliability, and deployment requirements. Classical, learned, and hybrid methods offer different ways to achieve that balance. Additional complexity should be justified by improvements demonstrated across the complete pipeline.

At Pixop, perceptual optimization is one of the areas we are actively investigating as part of our work on improving the balance between coding efficiency and visual quality.

Further Reading

[1] Y. Zhang, L. Zhu, G. Jiang, S. Kwong, and C.-C. J. Kuo, “A Survey on Perceptually Optimized Video Coding,” ACM Computing Surveys, vol. 55, no. 12, article 245, 2023. doi:10.1145/3571727. Preprint: arXiv:2112.12284.

[2] E. Vidal, N. Sturmel, C. Guillemot, P. Corlay, and F.-X. Coudoux, “New Adaptive Filters as Perceptual Preprocessing for Rate-Quality Performance Optimization of Video Coding,” Signal Processing: Image Communication, vol. 52, pp. 124–137, 2017. doi:10.1016/j.image.2016.12.003.

[3] E. Vidal, F.-X. Coudoux, P. Corlay, and C. Guillemot, “JND-Guided Perceptual Pre-filtering for HEVC Compression of UHDTV Video Contents,” ACIVS 2017, LNCS 10617, pp. 375–385, 2017. doi:10.1007/978-3-319-70353-4_32.

[4] B. Kang and W. Kim, “Human Perception-Oriented Enhancement and Smoothing for Perceptual Video Coding,” IEEE Transactions on Broadcasting, vol. 69, no. 3, 2023. doi:10.1109/TBC.2023.3291139.

[5] A. Chadha and Y. Andreopoulos, “Deep Perceptual Preprocessing for Video Coding,” Proc. IEEE/CVF CVPR, pp. 14852–14861, 2021.

[6] Y.-H. Sun, C. L.-H. Lee, and T.-S. Chang, “IQNet: Image Quality Assessment Guided Just Noticeable Difference Prefiltering for Versatile Video Coding,” IEEE Open Journal of Circuits and Systems, vol. 5, pp. 17–27, 2024. doi:10.1109/OJCAS.2023.3344094.

[7] S. Ki, J. Do, and M. Kim, “Learning-Based JND-Directed HDR Video Preprocessing for Perceptually Lossless Compression With HEVC,” IEEE Access, 2020. doi:10.1109/ACCESS.2020.3046194.

[8] Y. Pang, S. Zhao, J. Li, and L. Zhang, “Adaptive High-Frequency Preprocessing for Video Coding,” Proc. IEEE ICIP, pp. 151–156, 2025. doi:10.1109/ICIP55913.2025.11084703.

[9] M. U. K. Khan, A. Chadha, M. A. Anam, and Y. Andreopoulos, “Perceptual Video Compression with Neural Wrapping,” Proc. IEEE/CVF CVPR, pp. 17743–17754, 2025.

[10] C. He, Z. Dong, M. Li, Z. Hao, L. Huang, X. Zeng, and Y. Fan, “JND-Guided Light-Weight Neural Pre-Filter for Perceptual Image Coding,” arXiv:2510.10648, 2025.

From the Pixop team

All

Insights

Use Cases

News & Events

Load more

All

Insights

Use Cases

News & Events

Load more

All

Insights

Use Cases

News & Events

Load more

Ready to deliver consistent, premium video quality regardless of the source?

Ready to deliver consistent, premium video quality regardless of the source?

Talk to us about your workflow and we’ll show you how Pixop fits.

Talk to us about your workflow and we’ll show you how Pixop fits.