QP can be negative in HBD path.
- Change pu1_pic_qp and pu1_qp to WORD8 type
- Rename pu1_pic_qp and pu1_qp to pi1_pic_qp and pi1_qp
- Remove redundant typecasts
- Format new functions to be consistent with existing codebase conventions
- Move local variable declarations to top of function scope (C99 style)
- Avoid datatype declarations inside for loop headers and prefer WORD32
for loop indices without prefixes
- Remove commented code blocks
- Rename WORD32 arguments and header declarations to remove i4_ prefix
Added few members in codec_t and initialized them for 420P, 8bit path.
When HBD support is added these will be updated based on the parsed headers.
(cherry picked from commit f35c5b1971)
In ihevcd_copy_slice_hdr(), copying the previous slice header for dependent
slices copied i2_ctb_x and i2_ctb_y. If parsing was aborted or delayed,
the incomplete slice header retained the previous slice's coordinates,
causing concurrent worker threads in ihevcd_slice_hdr_update() to prematurely
switch to the invalid slice header. Reset i2_ctb_x and i2_ctb_y to -1
after copying.
In ihevcd_parse_slice_header(), initialize i2_ctb_x and i2_ctb_y to -1
immediately following memset so that slices rejected prior to coordinate
parsing do not retain valid-looking (0, 0) coordinates.
In ihevcd_parse_pic_init(), only skip reinitializing slice header 1 if
the first slice was missing in the current picture.
Test: ./hevc_dec_tests
Test: ./hevc_dec_fuzzer /work/misc/oss-fuzz/build/out/libhevc/crash-b74b07ae377d7450614e4534c86155d1d2999e13
TAG=agy
CONV=cfdfe257-7362-41c9-b5be-18261f1c6a3d
(cherry picked from commit 1713ca27ea)
When parsing PPS in ihevcd_parse_pps(), the scratch buffer
ps_codec->s_parse.ps_pps_base + MAX_PPS_CNT - 1 was not reset to zero,
unlike ihevcd_parse_sps(). In streams with concatenated sequences where
an earlier sequence set pps_extension_present_flag and range extension
flags, subsequent sequences without PPS extensions would retain the stale
flags. This caused the parser to attempt reading non-existent range
extension syntax, leading to bitstream parsing failure.
- In ihevcd_parse_pps(), reset the scratch PPS structure to zero upon
entry while preserving pi2_scaling_mat and ps_tile pointers.
- In both ihevcd_parse_pps() and ihevcd_parse_sps(), explicitly zero all
extension flags when pps_extension_present_flag / sps_extension_present_flag
is 0.
Test: ./hevc_dec_tests
TAG=agy
CONV=476d1054-10ff-4f19-88d5-a1c8dae7b57f
Currently, libhevc maintains two separate implementations for inverse
transform residual functions: ihevc_itrans_res_* and ihevc_hbd_itrans_res_*.
Since both 8-bit and HBD inverse transform residual routines output signed
16-bit residuals (WORD16 *pi2_dst), the only algorithmic difference between
8-bit and HBD variants is the 2nd stage shift:
Stage 1 shift: IT_SHIFT_STAGE_1 = 7 (fixed across all bit depths)
Stage 2 shift: shift = 20 - bit_depth
- 8-bit: 20 - 8 = 12 (IT_SHIFT_STAGE_2)
- 10-bit: 20 - 10 = 10
This commit cleans up the entire itrans_res family (4x4_ttype1, 4x4, dc, 8x8,
16x16, 32x32) as follows:
1. common/ihevc_itrans_res.h & common/ihevc_itrans_res.c:
- Add 'UWORD8 bit_depth' parameter to ihevc_itrans_res_* function
signatures and typedefs.
- Replace hardcoded 'shift = IT_SHIFT_STAGE_2;' with 'shift = 20 - bit_depth;'.
- Remove obsolete ihevc_hbd_itrans_res_* prototypes and typedefs.
2. common/ihevc_hbd_itrans_res.c:
- Delete file (~2,335 redundant lines removed).
3. Update tests/common/ihevc_itrans_res_test.cc to test bit depths 8 and 10
ihevcd_parse_sei() truncated the declared SEI payload size to the bytes
remaining with MIN() instead of rejecting it. The payload parsers do not
consult that size, so ihevcd_parse_mastering_disp_params_sei() read its
full 24-byte structure out of a one-byte payload and ran past the end of
the bitstream buffer.
(cherry picked from commit b21f809576)
When transquant bypass is enabled in HBD streams, decoded frames
mismatched the reference output due to two issues:
1. Intra boundary filtering:
In the 8-bit path, boundary filtering for horizontal (mode 10) and
vertical (mode 26) intra predictions is disabled when transquant bypass
is active with implicit RDPCM. However, the HBD reconstruction path was
passing raw u1_luma_pred_mode instead of disable_boundary_filter, and the
HBD filter functions ignored the mode parameter.
2. SAO Scratch Buffer Overflow:
During shifted CTB SAO processing, transquant bypass blocks are temporarily
backed up into pu1_tmp_buf_luma and restored afterwards to keep them lossless.
In HBD, each sample requires 2 bytes (UWORD16), but the static scratch buffer
was allocated for 8-bit samples (sizeof(UWORD8)).
Because the shifted CTB processing window spans beyond standard CTB boundaries
to handle delayed neighbor boundaries, the 16-bit luma backup exceeded the
allocated buffer capacity and overflowed into pu1_tmp_buf_chroma. Consequently,
chroma backup overwrote the luma backup data, resulting in corrupted samples
upon restoration.
This commit resolves the above issues with the following changes:
- common/ihevc_hbd_intra_pred_filters.c: Check 'disable_boundary_filter' to bypass
boundary filtering in ihevc_hbd_intra_pred_luma_horz and ihevc_hbd_intra_pred_luma_ver.
- decoder/ihevcd_iquant_itrans_recon_ctb.c: Pass disable_boundary_filter
to apf_hbd_intra_pred_luma for modes 10 and 26.
- decoder/ihevcd_api.c: Scale SAO temporary scratch buffer allocation
(pu1_tmp_buf_luma and pu1_tmp_buf_chroma) by sizeof(UWORD16) to prevent
buffer collision in HBD mode.
Previously, the reallocation check evaluated bytes_remaining (the total
unparsed input buffer passed to the decode call, which can contain
multiple frames). This caused unnecessary reallocations on large
multi-frame input chunks.
Update the check to look ahead for the next start code using
ihevcd_nal_search_start_code to find the exact current NAL unit size
(cur_nal_size). Trigger buffer reallocation only when cur_nal_size
exceeds ps_codec->u4_bitsbuf_size.
Additionally:
- Check the return status of ihevcd_reallocate_dynamic_bitstream_buf
and handle memory allocation failure cleanly by returning IV_FAIL with
IVD_MEM_ALLOC_FAILED.
- Directly assign bytes_remaining = cur_nal_size for emulation prevention
removal.
As per the HEVC specification, i1_log2_min_pcm_coding_block_size shall be in the range of:
Min(MinCbLog2SizeY, 5) to Min(CtbLog2SizeY, 5), inclusive.
Updated the validation to use Min(MinCbLog2SizeY, 5) as the lower bound.
When parsing SEI payloads before any valid SPS has been decoded,
ihevcd_parse_sei_payload scanned ps_codec->ps_sps_base using a loop
that could step past the array bounds and dereference an invalid pointer.
Initialize ps_sps to NULL, check each entry's validity within bounds,
and return safely if no valid SPS is present.
(cherry picked from commit d42200d2f2)
For High Bit Depth streams, pixel samples occupy 2 bytes (i4_pixel_size_y = 2).
In ihevcd_iquant_itrans_recon_ctb, the luma row offset in the destination buffer
for IPCM blocks was computed using `i * pic_strd` rather than `i * pic_strd *
i4_pixel_size_y`. This caused luma rows in IPCM blocks to overwrite previous rows
and corrupt reconstructed pixels.
Updated the row offset indexing to multiply by i4_pixel_size_y.
Test: ./hevcdec
- In 4:2:2 chroma format, transform blocks are partitioned into
two square sub-TUs (subtu_idx == 0 and subtu_idx == 1).
- Fix neighbor availability bitfield mapping for sub-TU 0 and sub-TU 1
based on trans_size (8, 16, 32).
- Add reference documentation from ITU-T H.265 Section 7.3.8.8 and
Section 8.4.4.2.2.
Test: ./hevcdec
Change-Id: I013e44e024258f2f84690fdc0109201129042155
- For CTB size 16 in 4:2:2, sub-block width is 0, causing the Top
and Current SAO blocks to be skipped.
- Fix Top-Right neighbor in Left block to read from picture buffer
when ctb_size == 16 and ctb_y != 0, while keeping line buffer for
larger CTBs (32, 64).
- Fix Bottom-Left neighbor in Left block to read from picture buffer
when ctb_size == 16 to avoid uninitialized backup buffer access.
- Generalize Top-Left block condition to (8 * v_samp_factor).
Test: ./hevcdec
Change-Id: I013e44e024258f2f84690fdc0109210129042245
This commit implements HBD inverse transform residual functions, which is a
prerequisite to enable 10-bit decoding with Cross Component Prediction (CCP).
When CCP is enabled, the inverse transform and prediction addition cannot be
performed in a single pass (`itrans_recon`). Instead, the decoder must first
compute the unclipped transform residuals (`itrans_res`), scale them for CCP,
and then add the prediction. For 10-bit bitstreams, the decoder was missing
HBD implementations for `itrans_res` and improperly falling back to 8-bit
implementations. This resulted in the use of an incorrect second-stage inverse
transform shift (hardcoded to 12 instead of `20 - bit_depth`), significantly
corrupting the residuals.
Key Changes:
1. Created `common/ihevc_hbd_itrans_res.c` providing mathematically compliant
HBD implementations of `itrans_res` for all Transform Unit sizes (4x4, 8x8,
16x16, 32x32, and DC) using intermediate `CLIP_S16` logic and dynamic
bit-depth shifting (`20 - bit_depth`).
2. Implemented `ihevc_hbd_chroma_recon_nxn_ccp` to properly scale 10-bit residuals
with the cross-component alpha and add the prediction.
3. Modified `ihevcd_iquant_itrans_resi_recon_tu_plane` to dynamically branch to
`ps_codec->apf_hbd_itrans_res` and `ihevc_hbd_chroma_recon_nxn_ccp` when
`pixel_size > 1`. This cleanly handles 10-bit CCP while keeping 8-bit paths
untouched.
Earlier condition in ihevcd_hbd_sao_shift_ctb was hardcoded
for 420 with luma ctb size as 16. Updated the condition to scale
for other chroma factors with vert and horz subsample factors.
Earlier condition was correctly hardcoded for 420 with
luma ctb size as 16. Updated the condition to scale for
other chroma factors with vert, and horz subsample factor.
Change-Id: I013e67e024258f2f84690fdcc01b4cd52f0fac52
(cherry picked from commit 7da5afcd72)
- Explicitly cast IHEVCD_SUCCESS to IHEVCD_ERROR_T in ihevcd_10bd_parse_sao to fix -Wimplicit-enum-enum-cast.
- Fix pointer comparison in ihevcd_parse_slice_data by replacing UWORD32 pointer cast with UWORD8* arithmetic to fix -Wpointer-to-int-cast.
With SIMD optimizations updated and enabled for 422 and 444 chroma formats,
the fallback override in ihevcd_init_function_ptr_rext_generic is no longer
necessary.
Test: ./build
(cherry picked from commit e3ce7e69cb)
Remove generic C overrides for chroma and luma intra prediction in ihevcd_init_function_ptr_rext_generic to allow architecture-specific assembly and SIMD intrinsics to be used during 422 and 444 decoding.
(cherry picked from commit efa8b27448)
This commit enables 10-bit decoding support for Implicit Residual DPCM
Key changes include:
- Removed the ASSERT(0) bypass limitation in ihevcd_iquant_itrans_resi_recon_tu_plane when i4_pixel_size_y > 1.
- Dynamically determining the bit_depth for luma/chroma plane.
- Branching the apf_recon calls to correctly use apf_hbd_recon for HBD reconstruction.
According to ITU-T H.265 Section 7.4.7.1, the initial SliceQpY value
shall be in the range of -QpBdOffsetY to +51, inclusive, where
QpBdOffsetY = 6 * (bit_depth_luma - 8).
Previously, the slice_qp_delta validity check in ihevcd_parse_slice_header
assumed an 8-bit lower bound of MIN_HEVC_QP (0), causing valid negative
slice QP values in high bit depth (10-bit/12-bit) streams to fail header
parsing with IHEVCD_INVALID_PARAMETER.
Updated the lower bound check to account for (6 * bit_depth_luma_minus8),
allowing valid HBD low QP bitstreams (such as QP = 0) to parse and decode
correctly.
Test: ./hevcdec
Apply Min(chroma_qp, 51) cap before adding QpBdOffsetC in HBD path for YUV422
and YUV444, matching the HEVC specification (Section 8.6.2).
Test: ./hevcdec
This commit introduces on-the-fly dynamic reallocation to safely decode
extreme edge-case bitstreams (e.g., random noise causing CABAC inflation)
that exceed standard MaxCPB size constraints.
It adds a dedicated `ihevcd_reallocate_dynamic_bitstream_buf` setup
function in `ihevcd_api.c`. If an incoming bitstream exceeds the currently
allocated internal buffer during parsing (`ihevcd_decode.c`), the decoder
will cleanly reallocate the buffer. This prevents truncation failures and
memory corruption for highly incompressible streams.
This commit implements the output formatting stage required to output
uncompressed 10-bit 444 frames to the application.
Changes include:
- Added `ihevcd_fmt_conv_hbd_444sp_to_444p` to handle de-interleaving of
16-bit U and V chroma channels into a planar format.
- Leveraged existing `ihevcd_fmt_conv_luma_copy` for 16-bit Luma copying
by doubling the width and stride parameters.
- Hooked the new conversion function into `ihevcd_fmt_conv` when processing
`IV_YUV_444P` formats with `pixel_size > 1`.