CN100477531C - Encoding method for compression encoding of multi-channel digital audio signal - Google Patents
Encoding method for compression encoding of multi-channel digital audio signal Download PDFInfo
- Publication number
- CN100477531C CN100477531C CNB200510095900XA CN200510095900A CN100477531C CN 100477531 C CN100477531 C CN 100477531C CN B200510095900X A CNB200510095900X A CN B200510095900XA CN 200510095900 A CN200510095900 A CN 200510095900A CN 100477531 C CN100477531 C CN 100477531C
- Authority
- CN
- China
- Prior art keywords
- sub
- band
- channel
- segment
- subband
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Expired - Lifetime
Links
- 238000000034 method Methods 0.000 title claims abstract description 74
- 230000005236 sound signal Effects 0.000 title claims abstract description 61
- 230000006835 compression Effects 0.000 title abstract description 27
- 238000007906 compression Methods 0.000 title abstract description 27
- 230000001052 transient effect Effects 0.000 claims abstract description 35
- 238000013139 quantization Methods 0.000 claims description 74
- 230000007774 longterm Effects 0.000 claims description 26
- 238000000354 decomposition reaction Methods 0.000 claims description 20
- 238000005070 sampling Methods 0.000 claims description 15
- 238000001514 detection method Methods 0.000 claims description 10
- 230000006870 function Effects 0.000 claims description 10
- 238000004458 analytical method Methods 0.000 claims description 6
- 238000001914 filtration Methods 0.000 claims 2
- 230000008569 process Effects 0.000 abstract description 18
- 230000005540 biological transmission Effects 0.000 abstract description 11
- 238000009826 distribution Methods 0.000 abstract description 7
- 238000005516 engineering process Methods 0.000 description 17
- 230000015572 biosynthetic process Effects 0.000 description 9
- 238000003786 synthesis reaction Methods 0.000 description 9
- 238000012545 processing Methods 0.000 description 6
- 230000001360 synchronised effect Effects 0.000 description 5
- 230000008901 benefit Effects 0.000 description 4
- 230000008859 change Effects 0.000 description 4
- 238000012546 transfer Methods 0.000 description 4
- 238000004364 calculation method Methods 0.000 description 3
- 238000012937 correction Methods 0.000 description 3
- 230000004044 response Effects 0.000 description 3
- 238000006243 chemical reaction Methods 0.000 description 2
- 230000007423 decrease Effects 0.000 description 2
- 238000010586 diagram Methods 0.000 description 2
- 230000000694 effects Effects 0.000 description 2
- 230000007704 transition Effects 0.000 description 2
- 238000013459 approach Methods 0.000 description 1
- 230000009286 beneficial effect Effects 0.000 description 1
- 230000001934 delay Effects 0.000 description 1
- 238000013461 design Methods 0.000 description 1
- 238000011161 development Methods 0.000 description 1
- 230000006872 improvement Effects 0.000 description 1
- 230000004807 localization Effects 0.000 description 1
- 238000004519 manufacturing process Methods 0.000 description 1
- 230000004048 modification Effects 0.000 description 1
- 238000012986 modification Methods 0.000 description 1
- 238000012827 research and development Methods 0.000 description 1
- 230000035945 sensitivity Effects 0.000 description 1
- 230000007480 spreading Effects 0.000 description 1
- 238000003892 spreading Methods 0.000 description 1
- 238000003860 storage Methods 0.000 description 1
Images
Landscapes
- Compression, Expansion, Code Conversion, And Decoders (AREA)
Abstract
Description
本申请是2002年8月21日提交、申请号为02130245.6的分案申请。This application is a divisional application submitted on August 21, 2002, with application number 02130245.6.
技术领域 technical field
本发明涉及数字音频信号的编码/解码设备及其方法,更确切地说,是关于对多声道数字音频信号进行压缩编码/解码的设备及方法。The present invention relates to a digital audio signal encoding/decoding device and its method, more precisely, to a multi-channel digital audio signal compression encoding/decoding device and method.
背景技术 Background technique
多声道(包括立体声)数字音频压缩编码技术已被广泛应用于VCD,SVCD,DVD,卫星电视,数字电视,和互联网(Internet)等领域中。它要解决的主要问题是用于表达多声道数字音频信号的码率很高,但可用于传播或储存它的信道容量却非常有限。例如,用PCM来表达48kHz采样率每个样本24比特的5.1声道的环绕声需要6912kbps(千比特/秒)的码率,而如数字电视之类的信道容量比较有限的应用可分配给数字音频信号的码率一般为384kbps,即使是DVD这样的信道容量比较宽松的应用可分配给数字音频信号的码率也一般为384kbps,768kbps,和1536kbps。在此,数字音频压缩编码技术需要提供高达18倍的压缩比。Multi-channel (including stereo) digital audio compression coding technology has been widely used in VCD, SVCD, DVD, satellite TV, digital TV, and Internet (Internet) and other fields. The main problem it needs to solve is that the code rate for expressing multi-channel digital audio signals is very high, but the channel capacity that can be used to propagate or store it is very limited. For example, using PCM to express 5.1-channel surround sound with 48kHz sampling rate and 24 bits per sample requires a code rate of 6912kbps (kilobits per second), and applications with limited channel capacity such as digital TV can be allocated to digital The code rate of the audio signal is generally 384kbps, and the code rates that can be assigned to the digital audio signal are generally 384kbps, 768kbps, and 1536kbps even for applications such as DVD with relatively loose channel capacity. Here, digital audio compression coding technology needs to provide up to 18 times the compression ratio.
数字音频压缩编码技术的研究开发可以追溯到70年代早期。经过三十年的发展,目前广泛采用的技术框架已基本定型为:帧长选择器,频率或子带分解器,暂态检测器,线性预测器,比特分配器,量化器,熵编码器,和多路复用器。The research and development of digital audio compression coding technology can be traced back to the early 1970s. After 30 years of development, the currently widely used technical framework has basically been finalized as follows: frame length selector, frequency or subband resolver, transient detector, linear predictor, bit allocator, quantizer, entropy encoder, and multiplexers.
例如,MPEG 2AAC[参考文献1]和MPEG 4AAC[参考文献2]技术把输入音频PCM信号流分成1024个样本一帧,然后对每帧信号作暂态检测。如果未发现本帧样本中有暂态响应,则(可选择地)作长期线性预测,然后作1024个子带的频率分解,再(可选择地)对每个子代信号作短期线性预测。如果发现有暂态响应,则进一步把本帧的1024个样本分成8个子帧,每帧128个样本,然后作128个子带的频率分解,并把暂态响应所在的那些子帧的位置传送给多路复用器。随后作基于人耳听觉模型的全局比特分配器,对子带信号作非线性标量量化,和对量化指数作哈夫曼(Huffman)编码。最后,多路复用器把以上各步骤所产生的辅助信息和表达各个子带样本的哈夫曼码打包成一个完整的以帧为单位的压缩码流。AAC的优点是压缩效率高。但编码器和解码器复杂,解码后的音频信号的音质不完全透明。For example, MPEG 2AAC [Reference 1] and MPEG 4AAC [Reference 2] techniques divide the input audio PCM signal stream into 1024 samples per frame, and then perform transient detection on each frame signal. If no transient response is found in the samples of this frame, then (optionally) perform long-term linear prediction, then perform frequency decomposition of 1024 sub-bands, and then (optionally) perform short-term linear prediction on each sub-generation signal. If a transient response is found, the 1024 samples of this frame are further divided into 8 subframes, each frame has 128 samples, and then the frequency decomposition of 128 subbands is performed, and the positions of those subframes where the transient response is located are sent to multiplexer. Then, a global bit allocator based on the human auditory model is implemented, nonlinear scalar quantization is performed on the subband signal, and Huffman (Huffman) encoding is performed on the quantization index. Finally, the multiplexer packs the auxiliary information generated by the above steps and the Huffman code expressing each sub-band sample into a complete frame-based compressed code stream. The advantage of AAC is its high compression efficiency. However, the encoder and decoder are complicated, and the sound quality of the decoded audio signal is not completely transparent.
再例如,DTS的多声道音频编码器[参考文献3,4和5]的帧长选择器可根据采样率和码率从256,512,1024,2048,和4096中选一个帧长,并按此帧长把输入音频PCM信号流分成帧。随后作32个子带的频率分解,再对每个子带信号作子带编码。子带编码包括暂态检测,线性预测,基于人耳听觉模型的全局比特分配器,标量/矢量量化,和哈夫曼(Huffman)编码。最后,多路复用器把以上各步骤所产生的辅助信息和表达各个子带样本的量化指数或哈夫曼码打包成一个完整的以帧为单位的压缩码流。DTS的优点是解码后的音频信号的音质好,在高码率(如1536kbps)时被很多人认为完全透明。但它的压缩效率不高。For another example, the frame length selector of DTS's multi-channel audio encoder [
随着数字电视近几年在欧州和北美的商业广播,多声道音频节目作为电视伴音的配送成为一个迫切需要解决的问题。这里涉及到的一个主要问题是目前的电视台和录音棚的设施仅仅支持立体声。把它们改成多声道意味着更换与音频相关的几乎全部设备。把多声道节目压缩到立体声能支持的码率即可避免这个问题。对多声道节目压缩后也有利于各个电视台和录音棚之间传输和分享节目。但压缩后的音频码流引入了帧的结构。如果音频帧的长度与视频帧的不等,在对视频码流在其帧的边界上作剪辑时就会切到音频帧的内部,从而破坏音频帧的结构,使解码器出错。另外,在多声道节目的制作和配送过程中往往需要对其进行多次的编码和解码(纵列编码Tandem Coding)操作。这要求压缩技术必须能经得起至少十次以上的纵列编码而听不到失真。With the commercial broadcast of digital TV in Europe and North America in recent years, the distribution of multi-channel audio programs as TV audio has become an urgent problem to be solved. One of the major issues involved here is that current television and recording studio facilities only support stereo. Changing them to multi-channel means replacing almost all equipment related to audio. This problem can be avoided by compressing multi-channel programs to a bit rate that stereo can support. Compressing multi-channel programs is also beneficial to the transmission and sharing of programs between various TV stations and recording studios. But the compressed audio stream introduces a frame structure. If the length of the audio frame is not equal to that of the video frame, when the video stream is clipped on the frame boundary, it will be cut into the audio frame, thereby destroying the structure of the audio frame and causing a decoder error. In addition, in the production and distribution process of multi-channel programs, it is often necessary to perform multiple encoding and decoding (Tandem Coding) operations. This requires that the compression technology must be able to withstand at least ten times of column coding without hearing distortion.
Dolby E是一个专为以上应用而设计的音频压缩编码技术[参考文献6]。它的帧长度固定为1792,但它用采样率变换的方法来使一帧Dolby E的数据流所占的时间与各种通用的视频帧频率(NTSC,PAL,和电影)的帧长度相等以达到能与它们同步剪辑的目的。同时,它又用高码率来确保能经得起十次以上的纵列编码而听不到失真。但是,Dolby E的压缩效率不高,不适合作为把多声道节目传输到最终用户(如电视机)的压缩技术。因此,电视台在用Dolby E制作好节目后还得解码成PCM,然后再编码成AC-3[参考文献7]或MPEG之类的高压缩效率的编码技术的码流后才能发射出去。图1示出采用Dolby E作节目配送的压缩编码技术,AC-3作节目传输的压缩编码技术的电视台配送和传输音频节目的过程。从中可以看出,这个电视台法方案存在以下几个困难:1)音频信号的失真大:Dolby E本身的采样率转换引入失真,从Dolby E格式的码流到AC-3格式的码流的转移编码(Transcoding)又引入新的失真。2)已发射过的节目很难再用:如图1所示,如果要再用已发射过的节目,它必须被解码成PCM再经Dolby E编码后才能与其他节目(如广告等)切换。在发射时还得重新经过从Dolby E解码到AC-3编码的转移编码过程。由于作最终传输的码率一般不高,已发射过的节目在经过以上这一串(AC-3解码->Dolby E编码->Dolby E解码->AC-3编码)转移编码后的音质很难或无法保证。3)由于Dolby E的输入和输出的采样率没有简单的倍频关系,其编码和解码器都很复杂昂贵。Dolby E is an audio compression coding technology designed for the above applications [Ref 6]. Its frame length is fixed at 1792, but it uses the method of sampling rate conversion to make the time occupied by a frame of Dolby E data stream equal to the frame length of various common video frame frequencies (NTSC, PAL, and film). To achieve the purpose of editing synchronously with them. At the same time, it uses a high bit rate to ensure that it can withstand more than ten times of tandem encoding without hearing distortion. However, Dolby E's compression efficiency is not high, and it is not suitable as a compression technology for transmitting multi-channel programs to end users (such as televisions). Therefore, after the TV station uses Dolby E to produce a good program, it has to be decoded into PCM, and then encoded into a code stream of high-compression-efficiency encoding technology such as AC-3 [Reference 7] or MPEG before it can be transmitted. Figure 1 shows the process of distributing and transmitting audio programs by TV stations using Dolby E as the compression coding technology for program distribution and AC-3 as the compression coding technology for program transmission. It can be seen from this that there are several difficulties in this TV station method: 1) The distortion of the audio signal is large: the sampling rate conversion of Dolby E itself introduces distortion, and the transfer of the code stream from the Dolby E format to the AC-3 format Coding (Transcoding) introduces new distortion. 2) The programs that have been transmitted are difficult to reuse: as shown in Figure 1, if the programs that have been transmitted are to be reused, they must be decoded into PCM and encoded by Dolby E before they can be switched with other programs (such as advertisements, etc.) . It has to go through the transfer encoding process from Dolby E decoding to AC-3 encoding again when transmitting. Since the code rate for the final transmission is generally not high, the sound quality of the programs that have been transmitted after the above series of transfer encoding (AC-3 decoding -> Dolby E encoding -> Dolby E decoding -> AC-3 encoding) is very low. Difficult or impossible to guarantee. 3) Since the sampling rate of Dolby E's input and output has no simple frequency multiplication relationship, its encoding and decoding are very complicated and expensive.
发明内容 Contents of the invention
本发明的第一方面,提供一种高效高保真的对多声道(包括单声道)音频信号进行压缩编码的编码器及其编码方法。当该音频信号作为视频信号的伴音时,该方法既能满足配送多声道数字音频节目的要求,也能满足以中低码率传播多声道数字音频节目的要求(压缩效率高)。也即,它实现了Dolby E和其它传输压缩编码技术如AC-3加起来的功能。The first aspect of the present invention provides a high-efficiency and high-fidelity encoder for compressing and encoding multi-channel (including mono) audio signals and its encoding method. When the audio signal is used as the accompanying sound of the video signal, the method can not only meet the requirements of distributing multi-channel digital audio programs, but also meet the requirements of transmitting multi-channel digital audio programs with medium and low bit rates (high compression efficiency). That is, it implements the combined functions of Dolby E and other transmission compression coding technologies such as AC-3.
本发明在该方面的编码器包括:1)帧长选择器,用于根据音频信号的采样率,码率,和视频帧频率(当多声道音频信号作为视频信号的伴音时)选择音频帧长;2)子带分解滤波器组,用于将一帧一帧输入的音频信号分解成多个子带信号;3)暂态检测器,用于将输入的子带信号分成暂态段与稳态段;4)比特分配器,用于将由目标码率所决定的比特资源分配到各个子带段;5)子带量化器,用于对所述的子带信号以段为单位量化;6)多路复用器(MUX),用于将子带的量化指数以及相关的辅助信息多路复用打包成一个以帧为单位的完整的码流。The encoder in this aspect of the present invention includes: 1) frame length selector, for selecting the audio frame according to the sampling rate of the audio signal, the code rate, and the video frame frequency (when the multi-channel audio signal is used as the accompanying sound of the video signal) long; 2) sub-band decomposition filter bank, used to decompose the input audio signal frame by frame into a plurality of sub-band signals; 3) transient detector, used to divide the input sub-band signal into transient segment and steady state segment 4) bit allocator, for distributing bit resources determined by the target code rate to each sub-band segment; 5) sub-band quantizer, for quantizing the sub-band signal in units of segments; 6 ) multiplexer (MUX), used to multiplex and package the sub-band quantization indices and related auxiliary information into a complete code stream in units of frames.
本发明在该方面的编码方法包括:1)根据音频信号的采样率,码率,和视频帧频率(当多声道音频信号作为视频信号的伴音时)选择音频帧长;2)通过子带分解滤波器组将一帧一帧输入的音频信号分解成多个子带信号;3)将各个子带信号分成暂态段与稳态段;4)将由目标码率所决定的比特资源分配到各个子带段;5)对所述的子带信号以段为单位量化;6)将子带信号的量化指数以及相关的辅助信息多路复用打包成一个以帧为单位的完整的码流。The encoding method of the present invention in this aspect comprises: 1) according to the sampling rate of audio signal, code rate, and video frame frequency (when multi-channel audio signal is used as the accompanying sound of video signal) select audio frame length; 2) by subband The analysis filter bank decomposes the input audio signal frame by frame into multiple sub-band signals; 3) divides each sub-band signal into a transient segment and a steady-state segment; 4) allocates bit resources determined by the target code rate to each Sub-band segment; 5) Quantize the sub-band signal in units of segments; 6) Multiplex and package the quantization index of the sub-band signal and related auxiliary information into a complete code stream in units of frames.
本发明的第二方面,提供一种对由上述编码器按上述编码方法编码形成的音频码流进行解码的解码器及其解码方法。其中该解码器包括:1)多路分解器(DEMUX),用于从上述编码的码流中多路分解出子带信号的量化指数以及相关的如音频帧长,子带段边界,和比特分配等辅助信息;2)子带逆量化器,用于依据相关的辅助信息以段为单位由子带信号的量化指数重建子带信号;3)子带合成滤波器组,用于由子带信号重建音频信号。The second aspect of the present invention provides a decoder for decoding an audio stream encoded by the above encoder according to the above encoding method and a decoding method thereof. Wherein the decoder includes: 1) demultiplexer (DEMUX), which is used to demultiplex the quantization index of the sub-band signal and related parameters such as audio frame length, sub-band segment boundary, and bit Auxiliary information such as allocation; 2) subband inverse quantizer, used to reconstruct subband signals from quantization indices of subband signals in units of segments according to related auxiliary information; 3) subband synthesis filter bank, used to reconstruct subband signals audio signal.
本发明在该方面的解码方法包括:1)从上述音频编码的码流中多路分解出子带信号的量化指数以及相关的辅助信息;2)依据相关的辅助信息以段为单位由子带信号的量化指数重建子带信号;3)用子带合成滤波器组由子带信号重建音频信号。The decoding method of the present invention in this aspect includes: 1) demultiplexing the quantization index of the sub-band signal and related auxiliary information from the code stream of the above-mentioned audio encoding; 3) Reconstruct the audio signal from the subband signal with the subband synthesis filter bank.
当多声道音频信号作为视频信号的伴音时,本发明选择的音频帧长与视频帧长有简单的倍数关系:或者视频帧长是音频帧长的正整数倍,或者音频帧长是视频帧长的整数倍。这样,本发明产生的码流可以与视频节目同步剪辑。When the multi-channel audio signal is used as the accompanying sound of the video signal, the audio frame length selected by the present invention has a simple multiple relationship with the video frame length: or the video frame length is a positive integer multiple of the audio frame length, or the audio frame length is the video frame length Integer multiples of long. In this way, the code stream generated by the present invention can be clipped synchronously with the video program.
更进一步,当多声道音频信号作为视频信号的伴音时,本发明选择的子带分解滤波器组的子带数是不同视频帧频(PAL和电影)所对应的不同音频帧长的公因子。这样,当视频节目的帧频发生变化时,本发明的编码器和解码器只须改变音频帧长与这个公因子的倍数,而不是子带数,即可维持音频帧长与视频帧长的前述倍数关系,以维持与视频节目的同步剪辑能力。由于子带数未变,编码器和解码器的各个延迟线也不用改变,因而编码器和解码器不用复位即可适应视频节目的帧频的变化.Furthermore, when the multi-channel audio signal is used as the accompanying sound of the video signal, the number of sub-bands of the sub-band decomposition filter bank selected by the present invention is a common factor of different audio frame lengths corresponding to different video frame rates (PAL and movies) . In this way, when the frame rate of the video program changes, the encoder and decoder of the present invention only need to change the multiple of the audio frame length and the common factor, rather than the number of subbands, to maintain the ratio between the audio frame length and the video frame length. The aforesaid multiple relationship is to maintain the synchronous editing ability with the video program. Since the number of sub-bands does not change, the delay lines of the encoder and decoder do not need to be changed, so the encoder and decoder can adapt to the change of the frame rate of the video program without resetting.
以上对子带数和帧长度的限制并不妨碍本发明有一个很短的帧长度以提供一个低延迟工作模式。The above limitations on the number of subbands and the frame length do not prevent the present invention from having a very short frame length to provide a low-delay working mode.
本发明的子带分解器与合成器所采用的滤波器较长以保证在对经本发明压缩编码后的数据流在帧的边界上作剪辑后能保持解码后的音频信号平滑过渡。The filters used by the sub-band decomposer and synthesizer of the present invention are relatively long to ensure smooth transition of the decoded audio signal after editing the data stream compressed and coded by the present invention at the frame boundary.
本发明的编码系统的暂态检测器把输入子带信号分割成暂态和平稳片段。它有很高的时间分辨率,以减小音频压缩编码常碰到的前向回声(Pre-echo)效应。本发明随后的各个部件对输入音频信号的处理都以片段为单位进行。The transient detector of the encoding system of the present invention splits the input subband signal into transient and stationary segments. It has a very high time resolution to reduce the forward echo (Pre-echo) effect often encountered in audio compression coding. Each subsequent component of the present invention processes the input audio signal in units of segments.
本发明利用跨声道的音频信号在同一子带(频率)上的相关性来作长期和短期预测。这样既可充分地利用音频信号的周期性及声道之间的相关性,又可采用低阶数的预测器以降低运算量。The present invention utilizes the correlation of audio signals across channels on the same subband (frequency) to make long-term and short-term predictions. In this way, the periodicity of the audio signal and the correlation between channels can be fully utilized, and a low-order predictor can be used to reduce the amount of computation.
本发明的比特分配器只非常有限地利用了人耳听觉模型以达到简化比特分配的目的。这样既可极大的减小比特分配的计算复杂度,还可仅仅用一个参数来表达比特分配的结果。解码器在接收到这个参数后即可根据它而很简单地准确重建编码器用的比特分配。这样就节省掉了其他技术用于传送比特分配的比特资源。由于这个节省,用这种方法为编码效率带来的改善很可能能与用很复杂的人耳听觉模型的媲美。The bit allocator of the present invention makes very limited use of the human auditory model for the purpose of simplifying bit allocation. In this way, the computational complexity of bit allocation can be greatly reduced, and only one parameter can be used to express the result of bit allocation. After receiving this parameter, the decoder can very simply reconstruct exactly the bit allocation used by the encoder from it. This saves the bit resource used by other techniques to transmit the bit allocation. Because of this savings, the improvement in coding efficiency with this approach is likely to be comparable to that achieved with a very complex model of the human ear.
在量化方面,本专利采用的策略是在量化级少时用矢量量化,量化级多时用标量量化,以达到优化压缩效率,解码复杂度和对记忆单元的需求的目的。In terms of quantization, the strategy adopted in this patent is to use vector quantization when there are few quantization levels, and use scalar quantization when there are many quantization levels, so as to achieve the purpose of optimizing compression efficiency, decoding complexity and the demand for memory units.
附图说明 Description of drawings
图1:采用Dolby的音频压缩技术的数字电视节目的配送和传输方案。Figure 1: Distribution and transmission scheme of digital TV programs using Dolby's audio compression technology.
图2:第一实施例的编码器。Figure 2: The encoder of the first embodiment.
图3:本发明暂态检测示意图。Figure 3: Schematic diagram of transient state detection in the present invention.
图4:第二实施例的编码器。Figure 4: Encoder of the second embodiment.
图5:人耳在安静环境下的听觉门限(Threshold in Quite)。Figure 5: Threshold in Quite of the human ear in a quiet environment.
图6:第三实施例的编码器。Figure 6: Encoder of a third embodiment.
图7:第四实施例的编码器。Figure 7: Encoder of a fourth embodiment.
图8:本声道线性预测与预测误差信号的量化。Figure 8: Quantization of linear prediction and prediction error signals for this channel.
图9:跨声道线性预测与预测误差信号的量化。Figure 9: Cross-channel linear prediction and quantization of the prediction error signal.
图10:本声道与跨声道线性预测并用,以及预测误差信号的量化。Figure 10: Use of local channel and cross-channel linear prediction, and quantization of the prediction error signal.
图11:编码流程图。Figure 11: Encoding flowchart.
图12:解码流程图。Figure 12: Decoding flowchart.
图13:解码器。Figure 13: Decoder.
图14:采用本声道线性预测时子带样本的重建过程。Figure 14: Reconstruction process of subband samples when using linear prediction for this channel.
图15:采用跨声道线性预测时子带样本的重建过程。Figure 15: The reconstruction process of subband samples when using cross-channel linear prediction.
图16:本声道和跨声道线性预测并用时子带样本的重建过程。Figure 16: Reconstruction process of sub-band samples with local and cross-channel linear prediction.
图17:采用本发明的编码器和解码器的数字电视节目的配送和传输示意图。Fig. 17: Schematic diagram of the distribution and transmission of digital TV programs using the encoder and decoder of the present invention.
具体实施方式 Detailed ways
本发明的编码及解码方法对各个声道的处理几乎是完全分离和相同的,因此以下对本发明的描述都基于一个声道,除非另外说明。The encoding and decoding method of the present invention is almost completely separated and identical to the processing of each channel, so the following description of the present invention is based on one channel, unless otherwise specified.
编码coding
[实施例1][Example 1]
如图2所示,本发明的编码器由帧长选择器1,子带分解滤波器组2、暂态检测器3、比特分配器4,子带量化器5,以及多路复用器6构成,用于对输入的音频信号进行压缩编码。该音频信号可以是数字电视的伴音信号。以下通过对各组件的详细描述来说明编码器的工作原理。As shown in Figure 2, the encoder of the present invention consists of a
帧长选择器1:Frame length selector 1:
帧长选择器1的作用是把输入的音频信号按照一定的帧长分成帧。至于帧长的选择,一般认为帧长越大,编码压缩效率越高,但编码延迟越大,编码器和解码器对记忆单元的需求量也越大。因此,码率高时一般选用小一些的帧长,码率低时一般选用大一些的帧长。由于编码延迟与帧长和采样率的乘积成正比,在选择帧长时一定得考虑采样率。The function of the
当音频信号是视频信号的伴音时,在选择帧长时除了要考虑以上因素外,还得考虑作为视频信号的伴音的一些特殊要求。这些特殊要求中最突出的,同时也是本发明要解决的问题之一,是被压缩后的音频码流能与视频信号作同步剪辑。本发明用“同步剪辑”来指被压缩后的音频码流与视频码流的合理剪辑点在时间上完全同步,以使得在它们的任何一个剪辑点同时对音视频码流作剪辑后,音视频码流都不出错,并且解码后的音视频信号在剪辑点都能平滑过渡。When the audio signal is the accompanying sound of the video signal, in addition to the above factors, some special requirements for the accompanying sound of the video signal must be considered when selecting the frame length. The most prominent of these special requirements, which is also one of the problems to be solved by the present invention, is that the compressed audio code stream can be edited synchronously with the video signal. The present invention uses "synchronous editing" to mean that the reasonable editing points of the compressed audio code stream and video code stream are fully synchronized in time, so that after editing the audio and video code stream at any of their editing points, the audio There is no error in the video code stream, and the decoded audio and video signals can transition smoothly at the editing point.
实现同步剪辑的基本要求是一帧音频码流所占的时间必须与一帧视频码流所占的时间有一个简单的倍数关系。否则,在视频帧的边界上剪辑视频码流时就会剪在音频帧的内部,破坏音频帧的结构,引起解码器出错。我们考虑两种情况。第一种为视频帧长是音频帧长的整数倍或相等的情况。此时可在每个视频帧的边界上作剪辑,而此剪辑点总会落在音频帧的边界上。The basic requirement for realizing synchronous editing is that the time occupied by one frame of audio stream must have a simple multiple relationship with the time occupied by one frame of video stream. Otherwise, when the video code stream is clipped on the boundary of the video frame, it will be clipped inside the audio frame, destroying the structure of the audio frame and causing a decoder error. We consider two cases. The first is the case where the video frame length is an integer multiple or equal to the audio frame length. At this point, clipping can be made on the boundary of each video frame, and the clip point will always fall on the boundary of the audio frame.
对于25帧/秒(PAL)的视频信号,其每帧所占的时间为0.04秒。相对于以48000kHz采样的音频信号,它相当于1920个样本。由于1920=3×5×27,这些因子的任何组合而形成的音频帧长都可与25帧/秒的视频信号做同步编辑。For a video signal of 25 frames per second (PAL), the time taken by each frame is 0.04 seconds. Relative to an audio signal sampled at 48000kHz, it is equivalent to 1920 samples. Since 1920=3×5×2 7 , the audio frame length formed by any combination of these factors can be edited synchronously with the 25 frame/second video signal.
对于24帧/秒(电影)的视频信号,与其对等的一帧音频信号(48000kHz采样率)的样本数为2000=24×53。同理,由这些因子组合而形成的音频帧长都可与视频信号做同步剪辑。For a video signal of 24 frames/second (movie), the number of samples of an equivalent frame of audio signal (48000kHz sampling rate) is 2000=2 4 ×5 3 . Similarly, the audio frame length formed by the combination of these factors can be edited synchronously with the video signal.
对音频帧长度选择的另一个限制是它必须是后面将要经历的子带分解滤波器组的子带数的整数倍。如果从1920和2000的最大公因子24×5中选一些组合作为子带数,我们则可在不变动子带数的条件下实现与24帧/秒和25帧/秒这两种视频信号的同步编辑。在子带分解滤波器组的描述中将对此做进一步的描述。Another restriction on the selection of the audio frame length is that it must be an integer multiple of the number of subbands of the subband decomposition filter bank to be experienced later. If some combinations are selected from the greatest
现在考虑第二种情况,即音频帧长是视频帧长的整数倍(如V倍)的情况。为了确保音频帧的完整性,必须每隔V个视频帧才有一个音频信号的剪辑点。被压缩过的视频码流往往要隔好几帧才会有一个剪辑点,而这个间隔还可能是动态的。假定这个动态间隔的最大值为W,则能同时满足视频和音频剪辑限制的最大剪辑间隔为WV个视频帧。这显然需要很大的存储量。但是,前面所描述的方法仍然适用。Now consider the second case, that is, the case where the audio frame length is an integer multiple (such as V times) of the video frame length. In order to ensure the integrity of the audio frame, there must be an audio signal clipping point every V video frames. The compressed video stream often has a clip point every several frames, and this interval may be dynamic. Assuming that the maximum value of this dynamic interval is W, then the maximum clip interval that can satisfy both video and audio clip constraints is WV video frames. This obviously requires a lot of storage. However, the methods described earlier still apply.
子带分解滤波器组2Subband
子带分解滤波器组2用于把由帧长选择器1传过来的音频信号分解成M个子带信号。在这里可以采用各种各样的滤波器组[参考文献8],包括但不限于完美重建和非完美重建滤波器组,正交和非正交滤波器组,树状滤波器组,余弦调制的滤波器组,子波包等。为了保证在对经本发明编码的码流进行剪辑后音频信号的平滑性,特别推荐使用重叠较长的滤波器组。The sub-band
由于运算量低和设计简单,余弦调制的滤波器组是本发明优选滤波器组。它的第k个子带滤波器为[参考文献8]:Due to the low calculation amount and simple design, the cosine modulated filter bank is the preferred filter bank of the present invention. Its kth subband filter is [Ref 8]:
其中n为样本指数,k为子带指数,M为子带数,p0(·)为原型滤波器,N为原型滤器的长度,
本发明对子带分解滤波器组2的子带数M有一定的限制,其原则是子带数M必须是对应不同的视频帧频率由帧长选择器1选择的音频帧长的一个公因子。也即,所有的可用的音频帧长都可由以下公式表达:The present invention has certain restriction to the sub-band number M of sub-band
Fsize=k*M (2)Fsize=k*M (2)
其中Fsize为音频帧长,k为一个正整数。这样,当视频的帧频发生变化时,帧长选择器1只需选择另一个k,就可在保持子带数M不变的条件下,选择一个与新的视频帧频对应的音频帧长。保持子带数不变就意味着编码器和解码器的各种延迟线不变。当视频信号的帧频率变化时,编码器和解码器都不需复位,只用对上式中的k作相应调整即能平稳地适应这个变化。Among them, Fsize is the audio frame length, and k is a positive integer. In this way, when the frame rate of the video changes, the
例如,在描述帧长选择器1时已提到,对于以48kHz采样的音频信号,2000可作为相对于24帧/秒的视频信号的音频帧长,1920可作为相对于25帧/秒的视频信号的音频帧长。如果选择子带数M为2000和1920的最大公因子80,则只须选择k=25和24即可在保持子带数不变的条件下分别实现2000和1920的音频帧长。当然,其它公因子,如40和20等,都可达到同样的效果。表1列出了本发明优选的音频帧长及滤波器组的子带数:For example, when describing the
表1Table 1
上表中的短帧长,如200,240,400,480等,很显然为本音频压缩方法提供了低延迟模式。因此,本发明对音频帧长及子带数的限制并不防碍本发明拥有低延迟模式。The short frame lengths in the above table, such as 200, 240, 400, 480, etc., obviously provide low-latency modes for this audio compression method. Therefore, the limitations of the present invention on the audio frame length and the number of subbands do not prevent the present invention from having a low-latency mode.
暂态检测器3
从子带分解滤波器组2输出的子带信号随即进入暂态检测器3。暂态检测器3依据一定的检测尺度来分析每帧子带信号的暂态情况,然后把一帧子带信号进一步分割成暂态段和稳态段,并输出每段的位置信息。需要特别强调的是,这些段位置自适应地随暂态情况而变化。本发明对子带信号的后续处理以段(以下称子带段)为基本单位。The sub-band signal output from the sub-band
以图3所示的子带信号为例,检测器3会判断A为稳态段,B为暂态段,C为稳态段,并输出每段的位置信息如段长30,20,和50。Taking the sub-band signal shown in Figure 3 as an example, the
用于暂态检测的尺度包括但不限于子带信号的能量,能量的对数,能量熵等。检测技术可以是简单的门限检测,也可以是一些很复杂的技术,包括但不限于著名的K-Means算法[参考文献8]。The scales used for transient detection include but not limited to the energy of the subband signal, the logarithm of the energy, the energy entropy, and so on. The detection technique can be simple threshold detection, or some very complex techniques, including but not limited to the well-known K-Means algorithm [Reference 8].
比特分配器4:Bit allocator 4:
比特分配器4可以用本技术领域常用的方法依据信噪比(SNR)或信号对遮盖比(SMR)为每个声道的每个子带段分配比特[参考文献1--9]。比特分配的结果为用于量化每个声道的每个子带段的每个子带样本的比特数。此结果被送给子带量化器5和多路复用器6。The bit allocator 4 can allocate bits for each subband segment of each channel according to the signal-to-noise ratio (SNR) or the signal-to-coverage ratio (SMR) by a method commonly used in the art [references 1-9]. The result of the bit allocation is the number of bits per subband sample used to quantize each subband segment for each channel. The result is sent to subband quantizer 5 and multiplexer 6 .
子带量化器5Subband Quantizer 5
子带量化器5包括一组多个拥有不同量化方式和量化级数的量化器[参考文献9]。这组量化器的选择与配置对压缩效率影响很大。由于量化指数的概率分布不均匀(尤其是在量化级少时),故一般压缩技术多采用标量量化加熵编码(如Huffman码)的方法。但熵编码的解码方法复杂,运算量大而且不均匀,为解码器的商业实现带来了不少困难。The sub-band quantizer 5 includes a set of quantizers with different quantization methods and quantization levels [Reference 9]. The selection and configuration of this group of quantizers has a great influence on the compression efficiency. Because the probability distribution of the quantization index is uneven (especially when there are few quantization levels), the general compression technology mostly adopts the method of scalar quantization plus entropy coding (such as Huffman code). However, the decoding method of entropy coding is complicated, and the amount of calculation is large and uneven, which brings many difficulties to the commercial realization of the decoder.
为此,本发明优选地采用在量化级少时用矢量量化,量化级多时用标量量化。For this reason, the present invention preferably adopts vector quantization when there are few quantization levels, and scalar quantization when there are many quantization levels.
子带量化器5依据比特分配的结果来为每个子带段从上述量化器组中选取一个具体的量化器,并以此量化器量化该子带段内的每一个子带样本。The sub-band quantizer 5 selects a specific quantizer from the quantizer group for each sub-band segment according to the bit allocation result, and quantizes each sub-band sample in the sub-band segment with this quantizer.
对每个子带段内的每个子带样本的量化过程分以下四步(图2):The quantization process of each subband sample in each subband segment is divided into the following four steps (Fig. 2):
1)估计比例因子:比例因子估计器51可以用该子带段的所有样本的绝对值的最大值,该子带段的所有样本的方差,或其它变量作为比例因子。1) Estimated scale factor: The
2)量化比例因子:比例因子本身也需要量化以便传送给解码器。由于人耳对音量的敏感度随音量增大而减小,故对比例因子的量化应采用非线性方式,如对数量化。以方差为例,设第c个声道第k个子带的第d个段的方差为σ(c,k,d),则比例因子量化器52为该比例因子选的的量化指数为:2) Quantization of the scale factor: the scale factor itself also needs to be quantized in order to be transmitted to the decoder. Since the sensitivity of the human ear to the volume decreases as the volume increases, the quantization of the scale factor should adopt a nonlinear method, such as logarithmic quantization. Taking the variance as an example, if the variance of the dth segment of the kth subband of the cth sound channel is σ(c, k, d), then the quantization index selected by the
其中α为量化步长。where α is the quantization step size.
3)归一化子带样本:子带样本归一化器53用量化后重建的比例因子对该子带段内的所有样本归一化。3) Normalize sub-band samples: the
4)量化子带样本:量化器选取器54依据比特分配器4送来的比特数选取具体的量化器,然后子带样本量化器55用它对归一化后的子带样本进行量化以得到每个样本的量化指数。4) Quantize the subband samples: the
多路复用器(MUX)6Multiplexer (MUX) 6
多路复用器6把以上各个编码器部件所产生的以下信息打包在一起以形成一个完整的比特流或码流。The multiplexer 6 packs together the following information generated by the above encoder components to form a complete bit stream or code stream.
帧长选择器1选择的音频帧长。The audio frame length selected by
暂态检测器3输出的段位置信息。Segment position information output by
比特分配器4为每个子带段分配的比特数。Number of bits allocated by
子带量化器5产生的比例因子的量化指数和每个子带样本的量化指数。The subband quantizer 5 produces the quantization indices of the scale factors and the quantization indices of each subband sample.
多路复用器6还会把其它辅助信息打包。这些辅助信息包括但不限于输入音频信号的采样频率,音箱设置,纠错码,时间码等。Multiplexer 6 also packs other auxiliary information. Such auxiliary information includes but is not limited to the sampling frequency of the input audio signal, speaker settings, error correction code, time code, etc.
[实施例2][Example 2]
本发明的第二实施例的编码器如图4所示。其大部分组件与实施例1的相同,区别在于本实施例的独特的比特分配方法。具体地说,本实施例的比特分配不象实施例1那样依据信噪比(SNR)以或信号对遮盖比(SMR)来为每个子带段分配比特,而是依据比例因子量化器52输出的比例因子量化指数用以下公式来为每个子带段分配比特。The encoder of the second embodiment of the present invention is shown in FIG. 4 . Most of its components are the same as those in
b(c,k,d)=f(α·s(c,k,d))-θ(k)-β(3)b(c,k,d)=f(α·s(c,k,d))-θ(k)-β(3)
其中:in:
1)b(c,k,d)是分配给当前子带段的每个样本的比特数。1) b(c, k, d) is the number of bits per sample allocated to the current subband segment.
2)f(α·s(c,k,d))是一个严格单调递增函数。它可被优选地设为f(α·s(c,k,d))=[α·s(c,k,d)]q,其中0<q≤2。它还可被进一步优选地设为f(α·s(c,k,d))=α·s(c,k,d)。2) f(α·s(c, k, d)) is a strictly monotonically increasing function. It can preferably be set as f(α·s(c, k, d))=[α·s(c, k, d)] q , where 0<q≦2. It can further preferably be set to f(α·s(c, k, d))=α·s(c, k, d).
3)θ(k)可以优选地设为如图5所示的的人耳在安静环境下的听觉门限(Threshold in Quite)的近似曲线[参考文献1,2和10],更可优选地设为零以简化计算。3) θ(k) can preferably be set as the approximate curve of the hearing threshold (Threshold in Quite) of the human ear in a quiet environment as shown in Figure 5 [
4)β是比特分配调整因子。4) β is the bit allocation adjustment factor.
从式(3)可看出,在比特分配调整因子β确定的情况下,分配到每个子带段的每个样本的比特数完全取决于其比例因子的量化指数。It can be seen from formula (3) that when the bit allocation adjustment factor β is determined, the number of bits allocated to each sample of each subband completely depends on the quantization index of its scale factor.
很明显,β越小(3)式分配给各个子带段的比特越多;β越大(3)式分配给各子带段的比特越少。比特分配的任务是在各个子带段所占用的比特数的总和不超过给定的目标码率(比特率)所允许分配给每帧音频信号的总比特数的条件下找到β的最小值。Obviously, the smaller β is, the more bits are allocated to each sub-band segment by formula (3); the larger β is, the less bits are allocated to each sub-band segment by formula (3). The task of bit allocation is to find the minimum value of β under the condition that the sum of the bits occupied by each sub-band segment does not exceed the total number of bits allocated to each frame of audio signal for a given target code rate (bit rate).
比特分配可以是全局最优的,也即所有的声道公用一个比特分配调整因子。假定在一个给定的目标码率下所允许分配给每帧音频信号的总比特数为Total Bits(这是在减去传送各种各样的辅助信息所需的比特数后的总值,以后总是如此假定,不再另外说明),比特分配调整因子搜索器41须搜索不同的β值以找到一个满足如下条件的最小β值:The bit allocation can be globally optimal, that is, all channels share a bit allocation adjustment factor. Assume that the total number of bits allowed to be allocated to each frame of audio signal at a given target bit rate is Total Bits (this is the total value after subtracting the number of bits required to transmit various auxiliary information, later always so assumed, no further explanation), the bit allocation
其中n(c,k,d)为子带段的样本数。where n(c, k, d) is the number of samples of the subband segment.
由此求得的β就可用来按照式(3)为所有声道的每个子带段分配比特。同时,编码器只须把这个比特分配调整因子传给解码器,解码器也就能依据它和比例因子的量化指数按照式(3)为所有声道的每个子带段重建编码器用的比特分配。The obtained β can be used to allocate bits for each subband segment of all channels according to formula (3). At the same time, the encoder only needs to pass this bit allocation adjustment factor to the decoder, and the decoder can reconstruct the bit allocation used by the encoder for each sub-band segment of all channels according to formula (3) according to it and the quantization index of the scale factor .
比特分配也可以是局部分别最优的,如每个声道分别拥有一个单独的比特分配调整因子。假定在一个给定的目标码率下按照某种预先确定的方式分配给第c声道的每帧音频信号的总比特数为Total Bits(c)。比特分配调整因子搜索器41须搜索不同的β值以找到一个满足如下条件的最小β值:The bit allocation can also be locally optimal, eg each channel has an individual bit allocation adjustment factor. Assume that the total number of bits of each frame of audio signal allocated to the c-th channel in a predetermined way at a given target bit rate is Total Bits(c). The bit allocation
此时,编码器就得为每个声道传送一个比特分配调整因子给解码器。显然,这个方法可以很直接地推广于用其它其它形式来分享比特分配调整因子的情况。In this case, the encoder has to transmit a bit allocation adjustment factor for each channel to the decoder. Apparently, this method can be extended directly to the situation of sharing the adjustment factor of bit allocation in other forms.
综上所述,本实施例的比特分配程序如下:In summary, the bit allocation procedure of this embodiment is as follows:
1)各个声道的所有比例因子量化器52把每个子带段的比例因子量化指数送给比特分配调整因子搜索器41。1) All
2)比特分配调整因子搜索器41依据式(3)和(4)寻找到一个在给定码率下的全局最优的比特分配调整因子,并传给多路复用器6以及比特分配器42。或者,比特分配调整因子搜索器41依据式(3)和(5)为每个声道分别寻找到一个在给定码率下的对各个声道局部最优的比特分配调整因子,并把它们分别传给多路复用器6以及各自的比特分配器42。2) The bit allocation
3)比特分配器42按式(3)为每个子带段分配比特,并传给其对应的量化器选取器54。3) The bit allocator 42 allocates bits for each sub-band segment according to formula (3), and sends the bit to its
由上可见,本发明的比特分配只非常有限地利用了人耳听觉模型以达到简化比特分配的目的。这样既可极大的减小比特分配的计算复杂度,还可仅仅用一个比特分配调整因子来表达比特分配的结果。在编码后的码流中仅须包含这个比特分配调整因子。解码器在接收到这个参数后可根据它用式(3)很简单地准确重建编码器用的比特分配。这样就节省掉了其他技术用于传送分配到每个子带段的比特数所需的比特资源。这些节省下来的比特资源可用于传送子带样本的量化指数,因而可以进一步提高音质。It can be seen from the above that the bit allocation of the present invention only uses the human auditory model very limitedly to achieve the purpose of simplifying the bit allocation. In this way, the computational complexity of bit allocation can be greatly reduced, and only one bit allocation adjustment factor can be used to express the result of bit allocation. Only this bit allocation adjustment factor must be included in the coded code stream. After receiving this parameter, the decoder can easily and accurately reconstruct the bit allocation used by the encoder using equation (3). This saves the bit resources required by other techniques to convey the number of bits allocated to each sub-band segment. These saved bit resources can be used to transmit the quantization index of the sub-band samples, thereby further improving the sound quality.
[实施例3][Example 3]
本发明的第三实施例的编码器如图6所示。其大部分组件与其它实施例的相同,区别在于本实施例比其它实施例多了一个联合强度编码器7。联合强度编码的理论基础是人耳对声音的空间定位在高频(如高于7kHz时)主要依据声音的强度。如图6所示,在编码时该联合强度编码器7可以把由左右(或其它可联合)声道的暂态检测器3输出的的高频子带加起来,只传这个和子带(称为源声道的被联合编码的子带)的各个样本的量化指数,外加被联合声道的被联合子带的强度指数,以达到节省比特的目的。The encoder of the third embodiment of the present invention is shown in FIG. 6 . Most of its components are the same as those of other embodiments, the difference is that this embodiment has one more joint strength encoder 7 than other embodiments. The theoretical basis of joint intensity coding is that the human ear's spatial localization of sound at high frequencies (such as above 7 kHz) is mainly based on the intensity of the sound. As shown in Figure 6, the joint strength encoder 7 can add up the high-frequency subbands output by the
当用到联合强度编码时,对源声道的被联合编码的子带段的比特分配必须考虑到其它被联合声道编码的同一子带段的比特需求。假定源声道为c,其他被联合编码的声道的总集为J,则在计算对c声道的被联合编码的子带段的比特分配时应采用的比例因子为:When joint intensity coding is used, the allocation of bits to jointly coded subband segments of source channels must take into account the bit requirements of other jointly coded subband segments of the same subband. Assuming that the source channel is c, and the total set of other jointly coded channels is J, the scale factor that should be used when calculating the bit allocation of the jointly coded subband segments of channel c is:
联合强度编码开始启用的频率如果过低可能会引起空间定位变窄,因此,在本实施例中,仅在低码率时才引用联合强度编码。If the starting frequency of the joint strength coding is too low, the spatial positioning may be narrowed. Therefore, in this embodiment, the joint strength coding is only used when the code rate is low.
实施例4Example 4
本发明的第四实施例的编码器如图7所示。其大部分组件与其它实施例的相同,区别在于本实施例比其它实施例多了一个跨声道的长期和短期线性预测器8。对每一段子带信号,本发明搜索它与本子带的短期和长期自相关性,及其与其它声道同一子带的信号的协相关性,以找到一个使预测误差最小的线性预测器8。设x(c,k,n)为第c个声道的第k个子带的第n个样本,则基于第s声道的对第c声道的线性预测为An encoder of a fourth embodiment of the present invention is shown in FIG. 7 . Most of its components are the same as those of other embodiments, the difference is that this embodiment has one more cross-channel long-term and short-term linear predictor 8 than other embodiments. For each sub-band signal, the present invention searches for its short-term and long-term autocorrelation with the sub-band, and its co-correlation with signals of the same sub-band of other channels to find a linear predictor that minimizes the prediction error 8 . Let x(c, k, n) be the n-th sample of the k-th subband of the c-th channel, then the linear prediction of the c-th channel based on the s-th channel is
其中,a(m)和b(m)分别为短期和长期预测器的预测滤波器的系数,τ为长期预测滤波器的延迟。当s=c时,上述预测完全基于本声道,m0=1;当s≠c时,上述预测是跨声道的,m0=0。where a(m) and b(m) are the coefficients of the prediction filters of the short-term and long-term predictors, respectively, and τ is the delay of the long-term prediction filter. When s=c, the above prediction is completely based on the current channel, m 0 =1; when s≠c, the above prediction is cross-channel, m 0 =0.
用(7)式对每一个子带样本x(c,k,n)做出预测后,就可得到一个对应的预测误差:After making predictions for each subband sample x(c, k, n) using formula (7), a corresponding prediction error can be obtained:
编码器的任务是找到一组最佳的预测系数a(m),b(m),和延迟τ以使在该子带段的总预测误差如以下的均方误差最小The task of the encoder is to find an optimal set of prediction coefficients a(m), b(m), and delay τ to minimize the total prediction error in this subband segment as the following mean square error
如果子带信号x(c,k,n)的周期性较强,则预测误差的方差会比子带信号x(c,k,n)本身的方差小很多。这意味着线性预测的预测增益很高,预测很成功。此时就可用预测误差信号e(c,k,n)来取代x(c,k,n)送到子带量化器5。否则,就直接把子带信号x(c,k,n)送到子带量化器5。因此,线性预测器8的工作程序可归纳如下:If the periodicity of the subband signal x(c, k, n) is strong, the variance of the prediction error will be much smaller than the variance of the subband signal x(c, k, n) itself. This means that the prediction gain of linear prediction is high and the prediction is successful. At this time, the prediction error signal e(c, k, n) can be used to replace x(c, k, n) and sent to the subband quantizer 5 . Otherwise, the subband signal x(c, k, n) is directly sent to the subband quantizer 5 . Therefore, the working procedure of linear predictor 8 can be summarized as follows:
1)估计预测系数与预测增益。1) Estimate the prediction coefficient and prediction gain.
2)如果预测增益高,将本子带段采用线性预测的决定、预测系数和延迟送到多路复用器6。同时,按(8)式产生预测误差信号,并将它送给量化器5。2) If the prediction gain is high, send the decision of linear prediction, prediction coefficient and delay to the multiplexer 6 for this sub-band segment. At the same time, a prediction error signal is generated according to (8) and sent to the quantizer 5 .
3)如果预测增益不高,将本子带段不采用线性预测的决定送到多路复用器6。同时,将本子带段的样本x(c,k,n)送给量化器5。3) If the prediction gain is not high, send the decision of not using linear prediction to the multiplexer 6 for this sub-band segment. At the same time, the samples x(c, k, n) of this sub-band are sent to the quantizer 5 .
当用到本声道预测时,为了避免量化误差在解码时扩散,对预测误差信号的量化必须闭环进行[参考文献8,9和11],如图8所示。由于解码器只能得到由量化器5输出的预测误差的量化指数,解码器只能用逆量化器9来重建预测误差,然后把它与预测值相加来重建子带样本。而本声道预测器81也只能用这些重建的子带样本来预测未来的子带样本。也即,代入(7)式的子带样本实质上是重建的子带样本。When local channel prediction is used, the quantization of the prediction error signal must be performed in a closed loop [
类似地,在作跨声道预测时,为了避免量化误差在解码时扩散,x(s,k,n)也必须是已解码后重建的子带样本,如图9所示。注意,图中9的预测器82是跨声道预测器,因为其输入的子带样本来自于另一个声道。Similarly, when making cross-channel prediction, in order to avoid the spread of quantization errors during decoding, x(s, k, n) must also be reconstructed subband samples after decoding, as shown in FIG. 9 . Note that
跨声道预测器82和本声道预测器81可以同时并用,如图10所示。首先,计算跨声道预测的误差:The
然后,对此误差作本声道预测:Then, the channel prediction is made for this error:
当用到跨声道预测时,编码器一定要确保在解码时源声道已经被解码。也就是说编码器一定要确保对各个声道的解码必须因循一定的可实现顺序:解码的第一个声道一定只能用本声道预测,第二个声道只能用本声道预测或前面已解码的那个声道作跨声道预测。以此类推。When cross-channel prediction is used, the encoder must ensure that the source channel is already decoded when decoding. That is to say, the encoder must ensure that the decoding of each channel must follow a certain achievable order: the first channel of decoding must only be predicted by this channel, and the second channel can only be predicted by this channel Or the previously decoded channel for cross-channel prediction. and so on.
编码器可以搜索所有的可实现解码顺序以找到预测增益最大的解码顺序。编码器也可只作有限的搜索以得到次优解。The encoder can search all achievable decoding orders to find the decoding order with the largest prediction gain. The encoder can also only do a limited search for a suboptimal solution.
预测范例1:长期或短期预测Forecasting Example 1: Long-Term or Short-Term Forecasting
由于长期和短期预测并用时估计预测系数较难,故可采用或者长期或者短期预测的方式以减小复杂度:Since it is difficult to estimate the prediction coefficient when long-term and short-term forecasting are used together, either long-term or short-term forecasting can be used to reduce complexity:
此时,τ小时为短期预测,τ大时为长期预测。At this time, when τ is small, it is short-term forecast, and when τ is large, it is long-term forecast.
当以上公式用在跨声道预测和本声道预测同时并用的情况时,首先计算跨声道预测的残差:When the above formula is used in the case of both cross-channel prediction and local-channel prediction, first calculate the residual of cross-channel prediction:
其中,延迟τ1可取的最小值为零。然后,对此残差作本声道预测:Among them, the minimum value of the delay τ 1 can be zero. Then, make this channel prediction for this residual:
其中,延迟τ2可取的最小值为1。Among them, the minimum value of the delay τ 2 is 1.
预测范例2:有限搜索的跨声道预测Prediction Example 2: Cross-Channel Prediction with Limited Search
为了简化跨声道预测在搜索解码次序上的复杂度,本范例先找到头两个声道的最优次序。对随后的声道仅搜索本声道预测以及用头两个声道作源声道的跨声道预测。In order to simplify the complexity of searching for the decoding order of cross-channel prediction, this example first finds the optimal order of the first two channels. Only local-channel predictions and cross-channel predictions using the first two channels as source channels are searched for subsequent channels.
为了进一步减小复杂度,可不做任何次序搜索,总用第一声道作为所有其它声道的跨声道预测的源声道。In order to further reduce the complexity, no sequential search may be performed, and the first channel is always used as the source channel for cross-channel prediction of all other channels.
较佳地,本发明的编码器还可以在音频帧输入到子带分解滤波器组2之前对音频帧进行跨声道和/或本声道预测,即在帧长选择器1与子带分解滤波器组2之间进一步包括一个跨声道和/或本声道预测器,此后,将预测误差输入到子带分解滤波器组2,并按照上面所描述的类似的处理对预测误差编码。Preferably, the encoder of the present invention can also perform cross-channel and/or local-channel prediction on the audio frame before the audio frame is input to the subband
编码流程:Coding process:
前述的四个实施例都可单独成为一个完整的编码器。但如果把它们的所有功能都汇总起来组成一个编码器,则可达到最佳压缩效率。图11示出了把本发明的四个实施例的所有功能汇合在一起时的编码流程。其中,联合强度编码,跨声道长短期预测,和全局比特分配是独立可选的。当它们中的任何一个未被选用时,其在图11中仅起一个把送入的数据不作任何改动地传出去的功能。当然,如果图11中的比特分配器没有被选用,另一个在实施例1中讨论过的通用比特分配器必须被引入以实现同样的功能。下面结合图2,4,6,7,8,9和10描述本发明的编码步骤(参见图11)。The aforementioned four embodiments can all be a complete encoder. However, if all their functions are combined to form an encoder, the best compression efficiency can be achieved. Fig. 11 shows the coding process when all the functions of the four embodiments of the present invention are combined. Among them, joint intensity coding, cross-channel long-term and short-term prediction, and global bit allocation are independently optional. When any one of them is not selected, it only plays a function of sending out the data sent in without any modification in Fig. 11 . Of course, if the bit allocator in Fig. 11 is not selected, another general bit allocator discussed in
E1)帧长选择:帧长选择器1接收从每个声道输入的音频信号的样本,根据音频信号的采样率,目标码率,和视频帧频率(当多声道音频信号作为视频信号的伴音时)选择音频帧长,然后把帧长信息传送给编码器的其它组件。由于本发明的编码器及其方法是以帧为基本单位,编码器的所有组件,编码流程的每一步,都直接或间接地用到帧长。但为了描述的简洁明朗,图11没有把帧长信息的传送路径全部标出。在帧长确定后,帧长选择器1还按帧长把输入音频信号的样本分成帧,并一帧一帧地送入子带分解滤波器组2。本步骤不对输入音频信号本身作任何处理。E1) frame length selection:
E2)子带分解:子带分解滤波器组2把每一声道的音频信号分解为M个子带信号。E2) Subband decomposition: the subband
E3)暂态检测:暂态检测器3分析每个子带信号的暂态情况,并据此把它分成暂态段和稳态段。然后,把每一段的位置信息传送给编码器的其它组件。由于与步骤E1类似的原因,图11没有把暂/稳态段的位置信息的传送路径全部标出。本步骤不对子带信号本身作任何处理。E3) Transient state detection: the
E4)联合强度编码:联合强度编码器7在作完联合强度编码后,抛弃被联合子带的样本,只将其每一子带段的强度指数传送给比特分配(E7)器4和多路复用(E9)器6。本步骤是本发明在低码率时优选的。不采用本步骤不影响编码器或方法的完整性,只是编码效率会有下降。E4) Joint strength coding: after joint strength coding, the joint strength encoder 7 discards the samples of the joint subbands, and only transmits the strength index of each subband segment to the bit allocator (E7)
E5)跨声道长期和短期预测:跨声道长期和短期预测器8在作完线性预测后,将是否采用预测的决定传给多路复用(E9)器6。如果决定采用预测,还将预测滤波器的延迟和预测系数传给多路复用(E9)器6,并用预测误差信号来取代子带信号传给子带量化器5。本步骤是本发明优选的,不采用本步骤不影响编码器或方法的完整性,只是编码效率会有下降。E5) Cross-channel long-term and short-term prediction: After the cross-channel long-term and short-term predictor 8 completes the linear prediction, it sends the decision whether to use prediction to the multiplexer (E9) 6. If it is decided to use prediction, the delay of the prediction filter and the prediction coefficient are also passed to the multiplexer (E9) 6, and the prediction error signal is used instead of the subband signal to the subband quantizer 5. This step is preferred in the present invention, and not using this step will not affect the integrity of the encoder or the method, but the encoding efficiency will decrease.
为了便于描述,图11把子带量化器5的功能分为两步:比例因子(E6)和矢量/标量量化(E8).For ease of description, Figure 11 divides the function of the subband quantizer 5 into two steps: scale factor (E6) and vector/scalar quantization (E8).
E6)比例因子:子带量化器5从子带信号(如决定采用了线性预测,则为预测误差信号,以下统称子带信号)中以暂/稳态段为单位估计并量化比例因子。然后,把比例因子的量化指数传送给比特分配(E7)器4和多路复用(E9)器6。本步骤不对子带信号本身作任何处理。E6) Scale factor: the subband quantizer 5 estimates and quantizes the scale factor in units of temporary/steady-state segments from the subband signal (if linear prediction is adopted, it is the prediction error signal, hereinafter collectively referred to as the subband signal). Then, the quantization indices of the scale factors are sent to the bit allocator (E7) 4 and the multiplexer (E9) 6. This step does not perform any processing on the sub-band signal itself.
E7)比特分配:比特分配调整因子搜索器41依据由步骤E6输入的比例因子的量化指数,以及由E4输入的强度指数(如果联合强度编码被选用),寻找到最优的比特分配调整因子,并把它传给多路复用(E9)器6。然后,比特分配器42又依据(3)式把比特分配到每一个子带段,并把每一个子带段所分配到的比特数传给量化器选取器54以供矢量/标量量化(步骤E8)用。以上的比特分配方法是本发明优选的。如不优选此方法,必须采用一个其它的比特分配方法,以维持编码器的完整性。本步骤不对子带信号本身作任何处理。E7) Bit allocation: the bit allocation
E8)矢量/标量量化:量化器选取器54依据比特分配(E7)器4送来的每一个子带段所分配到的比特数为每一个子带段选取一个量化器,然后把它传送给子带样本量化器55。子带样本量化器55随后以子带段为单位量化每一个子带样本,并把其量化指数传给多路复用(E9)器6。E8) vector/scalar quantization:
E9)多路复用(MUX):多路复用器6把每一个子带样本的量化指数和以下的辅助信息打包(多路复用)成一个完整的音频帧并将其输出:音频帧长,暂/稳态段的位置,强度指数(如果联合强度编码被选用),是否采用预测的决定,预测滤波器的延迟和系数,比例因子的量化指数,和比特分配调整因子。多路复用器6还可打包(多路复用)输出其它一些辅助数据,如采样频率,音箱设置,纠错码,时间码等。E9) multiplexing (MUX): the multiplexer 6 packs (multiplexes) the quantization index of each subband sample and the following auxiliary information into a complete audio frame and outputs it: audio frame Length, location of transient/stationary segments, intensity index (if joint intensity coding is selected), decision whether to use prediction, delay and coefficient of prediction filter, quantization index of scale factor, and bit allocation adjustment factor. The multiplexer 6 can also package (multiplex) and output other auxiliary data, such as sampling frequency, speaker settings, error correction code, time code and so on.
解码decoding
本发明的解码器及其解码方法在本质上为编码器及其方法的逆过程在此,根据图12和13描述解码流程和解码器各部件。The decoder and its decoding method of the present invention are essentially the inverse process of the encoder and its method. Here, the decoding process and components of the decoder are described based on FIGS. 12 and 13 .
解码流程:Decoding process:
经过本发明的编码方法和编码器产生的码流须经过以下的主要步骤(图12)来解码重建多声道音频信号:The code stream that produces through encoding method of the present invention and encoder must go through the following main steps (Fig. 12) to decode and reconstruct the multi-channel audio signal:
D1)解包(DEMUX)辅助信息:多路复用解包器110解包出以下辅助信息:D1) Unpacking (DEMUX) auxiliary information: the multiplexing
·帧长。• Frame length.
·所有子带段(暂/稳态段)的位置。• Location of all subband segments (transient/stationary segments).
·比特分配调整因子。• Bit Allocation Adjustment Factor.
·每个子带段的比例因子量化指数。• Scalefactor quantization index for each subband segment.
·每个子带段是否采用跨声道长期和短期预测的决定;如果采用,进一步解包出预测滤波器的延迟和系数。• A decision whether to use cross-channel long-term and short-term prediction for each subband segment; if so, further unpacking out the delays and coefficients of the prediction filter.
·被联合强度编码的子带段的强度指数The strength index of the subband segments that are jointly strength coded
D2)比特分配:比特分配器42依据输入的比特分配调整因子和每个子带段的比例因子量化指数为每个子带段分配比特。此比特分配器与编码器用到的比特分配器42完全一样,故仍延用标号42。本比特分配方法是本发明优选的;如编码器采用别的比特分配方法,则必须越过不作本步骤。但步骤D 1通常必须增加一个解包项目:从输入码流中解包出比特分配。D2) Bit allocation: The bit allocator 42 allocates bits for each sub-band segment according to the input bit allocation adjustment factor and the scale factor quantization index of each sub-band segment. This bit allocator is exactly the same as the bit allocator 42 used by the encoder, so the
D3)解包子带样本的量化指数:多路复用解包器110依据步骤D2分配的比特数从输入码流中解包出每个子带样本的量化指数。D3) Unpacking the quantization index of the sub-band samples: the multiplexing
D4)逆量化重建子带样本:子带样本逆量化器120依据步骤D1解包出的比例因子量化指数和步骤D3解包出的子带样本的量化指数重建每个子带样本。D4) Inverse quantization and reconstruction of sub-band samples: the sub-band sample
D5)跨声道长期和短期预测:对每一个子带段,如果步骤D1解包出的预测决定为肯定,则跨声道长期和短期预测器130对本子带段的所有样本作预测。否则,跨声道长期和短期预测器130对本子带段的所有样本不作任何处理。本步骤是本发明优选的;如编码器未采用跨声道长期和短期预测,则必须越过不作本步骤。D5) Cross-channel long-term and short-term prediction: For each sub-band segment, if the prediction decision unwrapped in step D1 is positive, then the cross-channel long-term and short-
D6)联合强度解码:对每一个被联合强度编码的子带段,联合强度解码器140首先从源子带段拷贝子带样本到本子带段。然后,依据步骤D1解包出的强度指数重建强度比例因子,并用它修正拷贝到本子带段的样本值。本步骤是本发明在低码率时优选的;如编码器未采用联合强度编码,则必须越过不作本步骤。D6) Joint intensity decoding: For each subband segment that is jointly intensity encoded, the
D7)合成滤波器组:合成滤波器组150把子带样本合成重建成音频信号。D7) Synthesis filter bank: The
如果本发明的编码器在音频帧输入到子带分解滤波器组2之前对音频帧进行了跨声道和/或本声道预测,此时还得对由合成滤波器组150重建的信号作相应的预测以重建音频信号。If the encoder of the present invention has carried out cross-channel and/or local-channel prediction to the audio frame before the audio frame is input to the sub-band
解码器:decoder:
本发明的解码器如图13所示。以下描述各个部件:The decoder of the present invention is shown in FIG. 13 . The individual components are described below:
多路复用解包器(DEMUX)110:Demultiplexer (DEMUX) 110:
多路复用解包器110负责从压缩编码的码流中解包出解码步骤D1和D3中列出的数据。它也负责从压缩编码的码流中解包出其它辅助数据,如采样频率,音箱设置,纠错码,时间码等。The
比特分配器42:Bit allocator 42:
此比特分配器与编码器用到的比特分配器42完全一样,故仍延用标号42。此比特分配器的功能是把比特分配调整因子和每个子带段的比例因子量化指数代入(3)式以得出分配给每个子带段的每个样本的比特数。This bit allocator is exactly the same as the bit allocator 42 used by the encoder, so the
子带样本逆量化器120:Subband Sample Inverse Quantizer 120:
子带样本逆量化器120根据每个子带段的比特分配选取量化器。然后,以此量化器和该子带段的比例因子由子带样本的量化指数重建子带样本。The subband sample
注意,当某一个子带段采用了跨声道长期和短期预测,则子带样本逆量化器120重建的是该子带段的预测误差信号,而不是子带样本本身。Note that when a certain sub-band segment adopts cross-channel long-term and short-term prediction, the
跨声道长期和短期预测器130:Cross-channel long-term and short-term predictors 130:
对每一个子带段,如果步骤D1解包出的预测决定为肯定,则跨声道长期和短期预测器130对本子带段的所有样本作预测。否则,跨声道长期和短期预测器130对本子带段的所有样本不作任何处理。For each sub-band segment, if the prediction decision unwrapped in step D1 is positive, the cross-channel long-term and short-
作本声道预测的预测器如图14。其中的预测器81与编码器用到的预测器81完全一样,故仍延用标号81。The predictor for this channel prediction is shown in Figure 14. The
作跨声道预测的预测器如图15。其中的预测器82与编码器用到的预测器82完全一样,故仍延用标号82。The predictor for cross-channel prediction is shown in Figure 15. The
当跨声道和本声道预测同时并用时,先作本声道预测,然后再作跨声道预测,如图16所示。When cross-channel and local-channel prediction are used together, the local-channel prediction is performed first, and then cross-channel prediction is performed, as shown in FIG. 16 .
联合强度解码器140:Joint intensity decoder 140:
对每一个被联合强度编码的子带段,联合强度解码器140首先从源子带段拷贝子带样本到本子带段。然后,依据强度指数重建强度比例因子,并用它修正拷贝到本子带段的样本值。For each subband segment to be jointly intensity coded, the
子带合成滤波器组150:Subband synthesis filter bank 150:
子带合成滤波器组150把子带样本合成重建成音频信号。子带合成滤波器组150是编码器中的子带分解滤波器组2的逆,是同时设计的。也即,编码器中的子带分解滤波器组2确定后,子带合成滤波器组150也就完全确定了[参考文献8]。The
应用方案application solution
本发明涉及的是多声道数字音频压缩编码/解码技术,其优点包括压缩效率高,解码器简单,解码后的音频信号保真度高,以及能适用于高,中,和低码率的各种应用。因此,它完全适用于纯粹的音频应用,如数字音频广播等;其解码器完全可独立地安装在纯粹的音响设备中,如功放,随身听等。The present invention relates to multi-channel digital audio compression encoding/decoding technology, which has the advantages of high compression efficiency, simple decoder, high fidelity of the decoded audio signal, and can be applied to high, medium, and low bit rates various applications. Therefore, it is completely suitable for pure audio applications, such as digital audio broadcasting, etc.; its decoder can be installed independently in pure audio equipment, such as power amplifiers, walkmans, etc.
由于音频信号大多以视频信号的伴音的形式出现,由于本发明的编解码方法的压缩效率高,经过本发明编码的音频信号能与视频信号同步剪辑并能经受得起十次以上的纵列编码,因而在实际应用中具有如下的优点:Because the audio signal mostly appears in the form of accompanying sound of the video signal, and because the compression efficiency of the encoding and decoding method of the present invention is high, the audio signal encoded by the present invention can be edited synchronously with the video signal and can withstand more than ten tandem encodings , so it has the following advantages in practical application:
1)能同时满足节目的配送和传输的要求。1) It can meet the requirements of distribution and transmission of programs at the same time.
2)本发明的编码技术极大地简化了节目配送的环节和设备。以数字电视为例,图17示出了采用本技术的节目配送和传输的过程。很明显,它比Dolby的方案(图1)简单了许多。2) The encoding technology of the present invention greatly simplifies the link and equipment of program distribution. Taking digital TV as an example, Fig. 17 shows the process of program distribution and transmission using this technology. Obviously, it is much simpler than Dolby's solution (Fig. 1).
3)由于省去了多次的转移编码环节,本技术极大地提高了节目配送过程的保真度。3) This technology greatly improves the fidelity of the program distribution process because multiple transfer coding links are omitted.
4)由于省去了图1中的多个编码器和解码器,本技术也极大地降低了节目配送的成本。4) Since multiple encoders and decoders in FIG. 1 are omitted, this technology also greatly reduces the cost of program distribution.
参考文献references
[1]ISO/IEC 13818-7,1997.[1] ISO/IEC 13818-7, 1997.
[2]ISO/IEC 14496-3,1998.[2] ISO/IEC 14496-3, 1998.
[3]S.Smyth,M.Smyth,and W.P.Smith,“Multi-channelPredictive Subband Audio conder using Psychoacoustic AdaptiveBit Allocation In Frequency,Time,And Over The MultipleChannel s,”US Patent 5956674.[3] S.Smyth, M.Smyth, and W.P.Smith, "Multi-channel Predictive Subband Audio conder using Psychoacoustic AdaptiveBit Allocation In Frequency, Time, And Over The MultipleChannels," US Patent 5956674.
[4]M.Smyth,“An overview of the Coherent Acoustics codingsystem,”http://www.dtsonline.com/whitepaper.pdf,1999.[4] M. Smyth, "An overview of the Coherent Acoustics coding system," http://www.dtsonline.com/whitepaper.pdf , 1999.
[5]S.Smyth,W.P.Smith,M.Smyth,M.Yan,and T.Jung,“DTS Coherent Acoustics Delivering High Quality MultichannelSound to the Consume,”AES 100th Convention,1996.[5] S.Smyth, W.P.Smith, M.Smyth, M.Yan, and T.Jung, "DTS Coherent Acoustics Delivering High Quality MultichannelSound to the Consume," AES 100th Convention, 1996.
[6]L.Fielder and C.Todd,“The design of a video friendlyaudio coding system for distribution applications,”AES 17thInternaltional Conference,pp.86-92,1999.[6]L.Fielder and C.Todd, "The design of a video friendly audio coding system for distribution applications," AES 17th Internaltional Conference, pp.86-92, 1999.
[7]C.Todd,G.Davidson,M.Davis,L.Fielder,B.Link,and S.Vernon,“AC-3,Flexible perceptual coding for audiotransmission and storage,”96th AES Convention,Amsterdam,1994.[7] C.Todd, G.Davidson, M.Davis, L.Fielder, B.Link, and S.Vernon, "AC-3, Flexible perceptual coding for audiotransmission and storage," 96 th AES Convention, Amsterdam, 1994 .
[8]P.P.Vaidyanathan,“Multirate systems and filterbanks,”Prentice Hall,1993.[8] P.P. Vaidyanathan, "Multirate systems and filterbanks," Prentice Hall, 1993.
[9]A.Gersho and R.M.Gray,“Vector quantization andsignal compression,”Kluwer,1992.[9] A.Gersho and R.M.Gray, "Vector quantization and signal compression," Kluwer, 1992.
[10]B.C.J.Moore,“An introduction to the psychologyof hearing,”Academic Press,1997.[10] B.C.J.Moore, "An introduction to the psychology of hearing," Academic Press, 1997.
[11]A.M.Kondoz,“Digital Speech,”John Wiley&Sons,1994.[11] A.M. Kondoz, "Digital Speech," John Wiley & Sons, 1994.
Claims (14)
Applications Claiming Priority (1)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| CN02130245.6A CN1233163C (en) | 2002-08-21 | 2002-08-21 | Compression encoding and decoding apparatus for multi-channel digital audio signal and method thereof |
Related Parent Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| CN02130245.6A Division CN1233163C (en) | 2002-08-21 | 2002-08-21 | Compression encoding and decoding apparatus for multi-channel digital audio signal and method thereof |
Publications (2)
| Publication Number | Publication Date |
|---|---|
| CN1783727A CN1783727A (en) | 2006-06-07 |
| CN100477531C true CN100477531C (en) | 2009-04-08 |
Family
ID=34144427
Family Applications (11)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| CNB2005100983032A Expired - Lifetime CN100505554C (en) | 2002-08-21 | 2002-08-21 | Method for decoding and reconstructing a multi-channel audio signal from an encoded audio data stream |
| CN02130245.6A Expired - Lifetime CN1233163C (en) | 2002-08-21 | 2002-08-21 | Compression encoding and decoding apparatus for multi-channel digital audio signal and method thereof |
| CNB2005100987141A Expired - Lifetime CN100474780C (en) | 2002-08-21 | 2002-08-21 | Decoding method for decoding and re-establishing multiple audio track audio signal from audio data stream after coding |
| CNB200510095900XA Expired - Lifetime CN100477531C (en) | 2002-08-21 | 2002-08-21 | Encoding method for compression encoding of multi-channel digital audio signal |
| CNB2005100987086A Expired - Lifetime CN100481734C (en) | 2002-08-21 | 2002-08-21 | Decoder for decoding and re-establishing multiple acoustic track audio signal from audio data code stream |
| CNB2005100987103A Expired - Lifetime CN100481735C (en) | 2002-08-21 | 2002-08-21 | Decoding method for decoding and re-establishing multiple and track audio signal from audio data stream after coding |
| CNB2005100987122A Expired - Lifetime CN100481736C (en) | 2002-08-21 | 2002-08-21 | Coding method for compressing coding of multiple audio track digital audio signal |
| CNB2005100987137A Expired - Lifetime CN100533990C (en) | 2002-08-21 | 2002-08-21 | Encoder for compression encoding of multi-channel digital audio signals |
| CNB2005100987118A Expired - Lifetime CN100435485C (en) | 2002-08-21 | 2002-08-21 | Decoder for decoding and re-establishing multiple audio track andio signal from audio data code stream |
| CNB2005100983028A Expired - Lifetime CN100481733C (en) | 2002-08-21 | 2002-08-21 | Coder for compressing coding of multiple sound track digital audio signal |
| CNB2005100987090A Expired - Lifetime CN100452657C (en) | 2002-08-21 | 2002-08-21 | Coding method for compressing coding of multiple audio track audio signal |
Family Applications Before (3)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| CNB2005100983032A Expired - Lifetime CN100505554C (en) | 2002-08-21 | 2002-08-21 | Method for decoding and reconstructing a multi-channel audio signal from an encoded audio data stream |
| CN02130245.6A Expired - Lifetime CN1233163C (en) | 2002-08-21 | 2002-08-21 | Compression encoding and decoding apparatus for multi-channel digital audio signal and method thereof |
| CNB2005100987141A Expired - Lifetime CN100474780C (en) | 2002-08-21 | 2002-08-21 | Decoding method for decoding and re-establishing multiple audio track audio signal from audio data stream after coding |
Family Applications After (7)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| CNB2005100987086A Expired - Lifetime CN100481734C (en) | 2002-08-21 | 2002-08-21 | Decoder for decoding and re-establishing multiple acoustic track audio signal from audio data code stream |
| CNB2005100987103A Expired - Lifetime CN100481735C (en) | 2002-08-21 | 2002-08-21 | Decoding method for decoding and re-establishing multiple and track audio signal from audio data stream after coding |
| CNB2005100987122A Expired - Lifetime CN100481736C (en) | 2002-08-21 | 2002-08-21 | Coding method for compressing coding of multiple audio track digital audio signal |
| CNB2005100987137A Expired - Lifetime CN100533990C (en) | 2002-08-21 | 2002-08-21 | Encoder for compression encoding of multi-channel digital audio signals |
| CNB2005100987118A Expired - Lifetime CN100435485C (en) | 2002-08-21 | 2002-08-21 | Decoder for decoding and re-establishing multiple audio track andio signal from audio data code stream |
| CNB2005100983028A Expired - Lifetime CN100481733C (en) | 2002-08-21 | 2002-08-21 | Coder for compressing coding of multiple sound track digital audio signal |
| CNB2005100987090A Expired - Lifetime CN100452657C (en) | 2002-08-21 | 2002-08-21 | Coding method for compressing coding of multiple audio track audio signal |
Country Status (1)
| Country | Link |
|---|---|
| CN (11) | CN100505554C (en) |
Families Citing this family (30)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN100505554C (en) * | 2002-08-21 | 2009-06-24 | 广州广晟数码技术有限公司 | Method for decoding and reconstructing a multi-channel audio signal from an encoded audio data stream |
| BRPI0418665B1 (en) * | 2004-03-12 | 2018-08-28 | Nokia Corp | method and decoder for synthesizing a mono audio signal based on the available multichannel encoded audio signal, mobile terminal and encoding system |
| TW200952462A (en) * | 2004-06-02 | 2009-12-16 | Panasonic Corp | Seamless switching between random access units multiplexed in a multi angle view multimedia stream |
| CN101312041B (en) * | 2004-09-17 | 2011-05-11 | 广州广晟数码技术有限公司 | Apparatus and methods for multichannel digital audio coding |
| KR100682915B1 (en) * | 2005-01-13 | 2007-02-15 | 삼성전자주식회사 | Multi-channel signal encoding / decoding method and apparatus |
| US8102872B2 (en) * | 2005-02-01 | 2012-01-24 | Qualcomm Incorporated | Method for discontinuous transmission and accurate reproduction of background noise information |
| WO2006091139A1 (en) * | 2005-02-23 | 2006-08-31 | Telefonaktiebolaget Lm Ericsson (Publ) | Adaptive bit allocation for multi-channel audio encoding |
| US20070036228A1 (en) * | 2005-08-12 | 2007-02-15 | Via Technologies Inc. | Method and apparatus for audio encoding and decoding |
| CN1758772B (en) * | 2005-11-04 | 2010-05-05 | 无敌科技(西安)有限公司 | Method for synchronous playing video and audio of medium document and its system |
| CN101030379B (en) * | 2007-03-26 | 2011-10-12 | 北京中星微电子有限公司 | Method and apparatus for allocating digital voice-frequency signal bit |
| CN101277461B (en) * | 2007-03-27 | 2011-05-04 | 展讯通信(上海)有限公司 | Method and system for simultaneously transmitting high-low code flow in TD-SCDMA system |
| CN101067931B (en) * | 2007-05-10 | 2011-04-20 | 芯晟(北京)科技有限公司 | Efficient configurable frequency domain parameter stereo-sound and multi-sound channel coding and decoding method and system |
| CN101321293B (en) * | 2007-06-06 | 2010-08-18 | 中兴通讯股份有限公司 | A device and method for realizing multiplexing of multiple programs |
| CN101179716B (en) * | 2007-11-30 | 2011-12-07 | 华南理工大学 | Audio automatic gain control method for transmission data flow of compression field |
| CN101562015A (en) * | 2008-04-18 | 2009-10-21 | 华为技术有限公司 | Audio-frequency processing method and device |
| JP4702402B2 (en) * | 2008-06-05 | 2011-06-15 | ソニー株式会社 | Signal transmitting apparatus, signal transmitting method, signal receiving apparatus, and signal receiving method |
| CN102157151B (en) * | 2010-02-11 | 2012-10-03 | 华为技术有限公司 | A multi-channel signal encoding method, decoding method, device and system |
| CN102436819B (en) * | 2011-10-25 | 2013-02-13 | 杭州微纳科技有限公司 | Wireless audio compression and decompression methods, audio coder and audio decoder |
| US9559651B2 (en) * | 2013-03-29 | 2017-01-31 | Apple Inc. | Metadata for loudness and dynamic range control |
| RU2704733C1 (en) | 2016-01-22 | 2019-10-30 | Фраунхофер-Гезелльшафт Цур Фердерунг Дер Ангевандтен Форшунг Е.Ф. | Device and method of encoding or decoding a multichannel signal using a broadband alignment parameter and a plurality of narrowband alignment parameters |
| CN108496221B (en) * | 2016-01-26 | 2020-01-21 | 杜比实验室特许公司 | adaptive quantization |
| CN110009803A (en) * | 2017-12-26 | 2019-07-12 | 广州因文际会品牌策划有限公司 | A kind of computer spare parts and accessories selling system Internet-based |
| CN108550369B (en) * | 2018-04-14 | 2020-08-11 | 全景声科技南京有限公司 | Variable-length panoramic sound signal coding and decoding method |
| CN111462767B (en) * | 2020-04-10 | 2024-01-09 | 全景声科技南京有限公司 | Incremental coding method and device for audio signal |
| CN111681664A (en) * | 2020-07-24 | 2020-09-18 | 北京百瑞互联技术有限公司 | Method, system, storage medium and equipment for reducing audio coding rate |
| CN112151046B (en) * | 2020-09-25 | 2024-06-18 | 北京百瑞互联技术股份有限公司 | Method, device and medium for adaptively adjusting multi-channel transmission bit rate of LC3 encoder |
| CN115691514B (en) * | 2021-07-29 | 2026-01-02 | 华为技术有限公司 | A method and apparatus for encoding and decoding multi-channel signals |
| CN115881139B (en) * | 2021-09-29 | 2026-04-28 | 华为技术有限公司 | Encoding and decoding methods, apparatus, devices, storage media and computer programs |
| CN118942471A (en) * | 2022-06-15 | 2024-11-12 | 腾讯科技(深圳)有限公司 | Audio processing method, device, equipment, storage medium and computer program product |
| CN118016080B (en) * | 2024-04-09 | 2024-06-25 | 腾讯科技(深圳)有限公司 | Audio processing method, audio processor and related device |
Family Cites Families (5)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| US5388181A (en) * | 1990-05-29 | 1995-02-07 | Anderson; David J. | Digital audio compression system |
| DE69331166T2 (en) * | 1992-02-03 | 2002-08-22 | Koninklijke Philips Electronics N.V., Eindhoven | Transmission of digital broadband signals |
| KR960012475B1 (en) * | 1994-01-18 | 1996-09-20 | 대우전자 주식회사 | Digital audio coder of channel bit |
| US5956674A (en) * | 1995-12-01 | 1999-09-21 | Digital Theater Systems, Inc. | Multi-channel predictive subband audio coder using psychoacoustic adaptive bit allocation in frequency, time and over the multiple channels |
| CN100505554C (en) * | 2002-08-21 | 2009-06-24 | 广州广晟数码技术有限公司 | Method for decoding and reconstructing a multi-channel audio signal from an encoded audio data stream |
-
2002
- 2002-08-21 CN CNB2005100983032A patent/CN100505554C/en not_active Expired - Lifetime
- 2002-08-21 CN CN02130245.6A patent/CN1233163C/en not_active Expired - Lifetime
- 2002-08-21 CN CNB2005100987141A patent/CN100474780C/en not_active Expired - Lifetime
- 2002-08-21 CN CNB200510095900XA patent/CN100477531C/en not_active Expired - Lifetime
- 2002-08-21 CN CNB2005100987086A patent/CN100481734C/en not_active Expired - Lifetime
- 2002-08-21 CN CNB2005100987103A patent/CN100481735C/en not_active Expired - Lifetime
- 2002-08-21 CN CNB2005100987122A patent/CN100481736C/en not_active Expired - Lifetime
- 2002-08-21 CN CNB2005100987137A patent/CN100533990C/en not_active Expired - Lifetime
- 2002-08-21 CN CNB2005100987118A patent/CN100435485C/en not_active Expired - Lifetime
- 2002-08-21 CN CNB2005100983028A patent/CN100481733C/en not_active Expired - Lifetime
- 2002-08-21 CN CNB2005100987090A patent/CN100452657C/en not_active Expired - Lifetime
Also Published As
| Publication number | Publication date |
|---|---|
| CN1750402A (en) | 2006-03-22 |
| CN100481734C (en) | 2009-04-22 |
| CN1750403A (en) | 2006-03-22 |
| CN100505554C (en) | 2009-06-24 |
| CN1750407A (en) | 2006-03-22 |
| CN100474780C (en) | 2009-04-01 |
| CN100481736C (en) | 2009-04-22 |
| CN1750406A (en) | 2006-03-22 |
| CN1750408A (en) | 2006-03-22 |
| CN100481735C (en) | 2009-04-22 |
| CN1233163C (en) | 2005-12-21 |
| CN1750404A (en) | 2006-03-22 |
| CN100533990C (en) | 2009-08-26 |
| CN1750409A (en) | 2006-03-22 |
| CN1783727A (en) | 2006-06-07 |
| CN100481733C (en) | 2009-04-22 |
| CN100452657C (en) | 2009-01-14 |
| CN1756087A (en) | 2006-04-05 |
| CN1750405A (en) | 2006-03-22 |
| CN1477872A (en) | 2004-02-25 |
| CN100435485C (en) | 2008-11-19 |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| CN100477531C (en) | Encoding method for compression encoding of multi-channel digital audio signal | |
| US9478224B2 (en) | Audio processing system | |
| US7769584B2 (en) | Encoder, decoder, encoding method, and decoding method | |
| RU2439718C1 (en) | Method and device for sound signal processing | |
| US9275648B2 (en) | Method and apparatus for processing audio signal using spectral data of audio signal | |
| KR101221918B1 (en) | A method and an apparatus for processing a signal | |
| US8364471B2 (en) | Apparatus and method for processing a time domain audio signal with a noise filling flag | |
| US20080140393A1 (en) | Speech coding apparatus and method | |
| JP2015535620A (en) | Apparatus and method for encoding and decoding an encoded audio signal using temporal noise / patch shaping | |
| EP2133872B1 (en) | Encoding device and encoding method | |
| JPWO2007026763A1 (en) | Stereo encoding apparatus, stereo decoding apparatus, and stereo encoding method | |
| TW201603004A (en) | Method and apparatus for decoding a compressed HOA representation, and method and apparatus for encoding a compressed HOA representation | |
| WO2009048239A2 (en) | Encoding and decoding method using variable subband analysis and apparatus thereof | |
| US7835915B2 (en) | Scalable stereo audio coding/decoding method and apparatus | |
| US20110019829A1 (en) | Stereo signal converter, stereo signal reverse converter, and methods for both | |
| US7725324B2 (en) | Constrained filter encoding of polyphonic signals | |
| WO2006041055A1 (en) | Scalable encoder, scalable decoder, and scalable encoding method | |
| KR100750115B1 (en) | Audio signal encoding and decoding method and apparatus therefor | |
| US20080162148A1 (en) | Scalable Encoding Apparatus And Scalable Encoding Method | |
| CN105336334B (en) | Multi-channel sound signal coding method, decoding method and device | |
| CN1783726B (en) | Decoder for decoding and reestablishing multi-channel audio signal from audio data code stream | |
| EP1639580B1 (en) | Coding of multi-channel signals | |
| Cavagnolo et al. | Introduction to Digital Audio Compression | |
| Noll | Digital audio for multimedia | |
| You et al. | DRA Audio Coding Standard: An Overview |
Legal Events
| Date | Code | Title | Description |
|---|---|---|---|
| C06 | Publication | ||
| PB01 | Publication | ||
| C10 | Entry into substantive examination | ||
| SE01 | Entry into force of request for substantive examination | ||
| C14 | Grant of patent or utility model | ||
| GR01 | Patent grant | ||
| EE01 | Entry into force of recordation of patent licensing contract |
Application publication date: 20060607 Assignee: Shenzhen Guangsheng Digital Technology Co.,Ltd. Assignor: DIGITAL RISE TECHNOLOGY CO.,LTD. Contract record no.: 2010990000326 Denomination of invention: Coding method for compressing coding of multiple audio track digital audio signal Granted publication date: 20090408 License type: Common License Record date: 20100602 |
|
| LICC | Enforcement, change and cancellation of record of contracts on the licence for exploitation of a patent or utility model | ||
| CX01 | Expiry of patent term |
Granted publication date: 20090408 |
|
| CX01 | Expiry of patent term |

