WO2025234858A1 - Dispositif électronique et procédé de traitement de signal audio - Google Patents

Dispositif électronique et procédé de traitement de signal audio

Info

Publication number
WO2025234858A1
WO2025234858A1 PCT/KR2025/095221 KR2025095221W WO2025234858A1 WO 2025234858 A1 WO2025234858 A1 WO 2025234858A1 KR 2025095221 W KR2025095221 W KR 2025095221W WO 2025234858 A1 WO2025234858 A1 WO 2025234858A1
Authority
WO
WIPO (PCT)
Prior art keywords
channel
layout
audio signal
electronic device
audio signals
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Pending
Application number
PCT/KR2025/095221
Other languages
English (en)
Korean (ko)
Inventor
김경래
권용민
남우현
손윤재
정현권
황성희
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
Samsung Electronics Co Ltd
Original Assignee
Samsung Electronics Co Ltd
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Priority claimed from KR1020240194621A external-priority patent/KR20250160809A/ko
Application filed by Samsung Electronics Co Ltd filed Critical Samsung Electronics Co Ltd
Publication of WO2025234858A1 publication Critical patent/WO2025234858A1/fr
Pending legal-status Critical Current
Anticipated expiration legal-status Critical

Links

Classifications

    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06FELECTRIC DIGITAL DATA PROCESSING
    • G06F16/00Information retrieval; Database structures therefor; File system structures therefor
    • G06F16/60Information retrieval; Database structures therefor; File system structures therefor of audio data
    • G06F16/68Retrieval characterised by using metadata, e.g. metadata not derived from the content or metadata generated manually
    • GPHYSICS
    • G10MUSICAL INSTRUMENTS; ACOUSTICS
    • G10LSPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
    • G10L19/00Speech or audio signals analysis-synthesis techniques for redundancy reduction, e.g. in vocoders; Coding or decoding of speech or audio signals, using source filter models or psychoacoustic analysis
    • G10L19/008Multichannel audio signal coding or decoding using interchannel correlation to reduce redundancy, e.g. joint-stereo, intensity-coding or matrixing
    • GPHYSICS
    • G10MUSICAL INSTRUMENTS; ACOUSTICS
    • G10LSPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
    • G10L19/00Speech or audio signals analysis-synthesis techniques for redundancy reduction, e.g. in vocoders; Coding or decoding of speech or audio signals, using source filter models or psychoacoustic analysis
    • G10L19/04Speech or audio signals analysis-synthesis techniques for redundancy reduction, e.g. in vocoders; Coding or decoding of speech or audio signals, using source filter models or psychoacoustic analysis using predictive techniques
    • G10L19/16Vocoder architecture
    • HELECTRICITY
    • H04ELECTRIC COMMUNICATION TECHNIQUE
    • H04SSTEREOPHONIC SYSTEMS 
    • H04S3/00Systems employing more than two channels, e.g. quadraphonic

Definitions

  • An electronic device and method for processing an audio signal are disclosed. Specifically, an electronic device and method for processing an audio signal including object audio signals are disclosed.
  • Object audio signals are audio signals implemented to allow users to experience immersive sound, which is the sound of objects moving in three-dimensional space.
  • Object audio signals can be converted into multi-channel audio signals through rendering, so that they can be output to speaker systems with a multi-channel layout.
  • a method for an electronic device to process an audio signal may include a step of rendering a plurality of object audio signals to obtain metadata of the plurality of object audio signals.
  • the method may include a step of compressing the plurality of object audio signals, the metadata, and a bed channel audio signal to generate a bitstream.
  • the metadata may include, for each of a plurality of multi-channel layouts, channels on which the plurality of object audio signals are panned and gain values for each of the panned channels.
  • a method for processing an audio signal by an electronic device may include a step of decompressing a bitstream to obtain a plurality of object audio signals, metadata of the plurality of object audio signals, and a bed channel audio signal.
  • the method may include a step of identifying a target multi-channel layout that minimizes loss during channel audio rendering with a playback loudspeaker layout among a plurality of multi-channel layouts of the metadata.
  • the method may include a step of performing channel mixing on the plurality of object audio signals based on the metadata to obtain a first multi-channel audio signal of the target multi-channel layout.
  • the method may include a step of performing channel audio rendering on the first multi-channel audio signal to obtain a second multi-channel audio signal of the playback loudspeaker layout.
  • the metadata may include, for each of the plurality of multi-channel layouts, channels on which the plurality of object audio signals are panned and a gain value for each of the panned channels.
  • an electronic device for processing audio signals may include a memory storing one or more instructions and at least one processor for executing one or more instructions stored in the memory.
  • the electronic device may render a plurality of object audio signals by executing one or more instructions, thereby obtaining metadata of the plurality of object audio signals.
  • the electronic device may compress a plurality of object audio signals, metadata, and a bed channel audio signal, thereby generating a bitstream.
  • the metadata may include, for each of a plurality of multi-channel layouts, channels on which a plurality of object audio signals are panned and gain values for each of the panned channels.
  • an electronic device for processing an audio signal may include a memory storing one or more instructions and at least one processor for executing one or more instructions stored in the memory.
  • the electronic device may decompress a bitstream to obtain a plurality of object audio signals, metadata of the plurality of object audio signals, and a bed channel audio signal.
  • the electronic device may identify a target multi-channel layout that minimizes loss during channel audio rendering as a playback loudspeaker layout among a plurality of multi-channel layouts of metadata.
  • the electronic device may perform channel mixing on a plurality of object audio signals based on metadata to obtain a first multi-channel audio signal of the target multi-channel layout.
  • at least one processor may execute one or more instructions to cause the electronic device to perform channel audio rendering on a first multi-channel audio signal to obtain a second multi-channel audio signal of a playback loudspeaker layout.
  • the metadata may include, for each of a plurality of multi-channel layouts, channels on which a plurality of object audio signals are panned and gain values for each of the panned channels.
  • a computer-readable recording medium having recorded thereon a program for performing an operation of an electronic device, one of the methods described above and below.
  • FIG. 1 is a drawing schematically illustrating the operation of electronic devices according to one embodiment of the present disclosure.
  • FIG. 2 is a flowchart illustrating a method for encoding an audio signal by an electronic device according to one embodiment of the present disclosure.
  • FIG. 3 is a diagram for explaining an operation of an electronic device encoding an audio signal according to one embodiment of the present disclosure.
  • FIG. 4 is a diagram for explaining metadata of an object audio signal according to one embodiment of the present disclosure.
  • FIG. 5 is a diagram illustrating a method for an electronic device according to an embodiment of the present disclosure to decode an audio signal.
  • FIG. 6 is a diagram for explaining an operation of an electronic device decoding an audio signal according to one embodiment of the present disclosure.
  • FIG. 7 is a diagram illustrating channel mixing and channel audio rendering according to one embodiment of the present disclosure.
  • FIG. 8 is a diagram for explaining an operation of an electronic device according to an embodiment of the present disclosure to perform rendering of grouped object audio signals.
  • FIG. 9 is a schematic diagram of an electronic device for encoding an audio signal according to one embodiment of the present disclosure.
  • FIG. 10 is a schematic diagram of an electronic device for decoding an audio signal according to one embodiment of the present disclosure.
  • FIG. 11 is a detailed configuration diagram of an electronic device according to one embodiment of the present disclosure.
  • the expression “configured to” can be used interchangeably with, for example, “suitable for,” “having the capacity to,” “designed to,” “adapted to,” “made to,” or “capable of.”
  • the term “configured to” does not necessarily mean something is “specifically designed to” in hardware.
  • the expression “a system configured to” can mean that the system, in conjunction with other devices or components, is “capable of.”
  • the phrase “a processor configured to perform A, B, and C” can mean a dedicated processor for performing the operations (e.g., an embedded processor), or a generic-purpose processor (e.g., a CPU or application processor) that can perform the operations by executing one or more software programs stored in memory.
  • a processor is a component that controls a series of processes so that an electronic device operates according to the embodiments described below, and may be composed of one or more processors.
  • One or more processors included in the processor may be circuitry such as a System on Chip (SoC), an Integrated Circuit (IC), etc.
  • One or more processors included in the processor may be a general-purpose processor such as a Central Processing Unit (CPU), a Micro Processor Unit (MPU), an Application Processor (AP), a Digital Signal Processor (DSP), a graphics-only processor such as a Graphics Processing Unit (GPU), a Vision Processing Unit (VPU), an artificial intelligence-only processor such as a Neural Processing Unit (NPU), or a communication-only processor such as a Communication Processor (CP).
  • the artificial intelligence-only processor may be designed with a hardware structure specialized for processing a specific artificial intelligence model.
  • a processor in the present disclosure may include various processing circuits and/or multiple processors.
  • the term “processor” as used herein, including in the claims, may include various processing circuits, including at least one processor. At least one processor, and one or more processors may be configured to perform, individually and/or collectively, the various functions described herein in a distributed fashion.
  • “processor,” “at least one processor,” and “one or more processors” may be configured to perform multiple functions. However, these terms encompass, without limitation, situations where one processor performs some of the functions and other processor(s) perform other parts of the functions, and situations where a single processor may perform all of the functions.
  • the at least one processor may include a combination of processors that perform various of the functions described in a distributed manner. The at least one processor may execute program instructions to accomplish or perform the various functions.
  • the processor can write data to memory, read data stored in memory, and process data according to predefined operating rules or artificial intelligence models, particularly by executing a program or at least one instruction stored in memory. Accordingly, the processor can perform the operations described in the following embodiments, and operations described as performed by electronic devices or detailed components included in the electronic devices in the following embodiments can be considered to be performed by the processor, unless otherwise specified.
  • a single processor or a combination of processors may include circuitry that performs processing, such as an Application Processor (AP), a Communication Processor (CP), a Graphical Processing Unit (GPU), a Neural Processing Unit (NPU), a Microprocessor Unit (MPU), a System on Chip (SoC), or an Integrated Chip (IC).
  • AP Application Processor
  • CP Communication Processor
  • GPU Graphical Processing Unit
  • NPU Neural Processing Unit
  • MPU Microprocessor Unit
  • SoC System on Chip
  • IC Integrated Chip
  • each block of the flowchart drawings and combinations of the flowchart drawings can be performed by computer program instructions.
  • These computer program instructions can be installed in a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing equipment, such that the instructions, when executed by the processor of the computer or other programmable data processing equipment, create a means for performing the functions described in the flowchart block(s).
  • These computer program instructions can also be stored in a computer-available or computer-readable memory that can direct a computer or other programmable data processing equipment to perform the functions in a specific manner, such that the instructions stored in the computer-available or computer-readable memory can produce an article of manufacture that includes instruction means for performing the functions described in the flowchart block(s).
  • the computer program instructions may be installed on a computer or other programmable data processing device, a series of operational steps may be performed on the computer or other programmable data processing device to create a computer-executable process, and the instructions that cause the computer or other programmable data processing device to perform the steps for performing the functions described in the flowchart block(s) may also provide steps for performing the functions described in the flowchart block(s).
  • each block may represent a module, segment, or portion of code that contains one or more executable instructions for performing a specific logical function(s).
  • the functions described in the blocks may occur out of order. For example, two blocks depicted in succession may actually be executed substantially concurrently, or the blocks may sometimes be executed in reverse order, depending on their respective functions.
  • the term "electronic device” may refer to any device that processes an audio signal.
  • the “electronic device” may perform at least one of the encoding method or decoding method for an audio signal described in the present disclosure.
  • an 'audio signal' may mean a signal including information related to audio.
  • the 'audio signal' may be a 'channel-based audio signal' or an 'object-based audio signal', or may include both a 'channel-based audio signal' and an 'object-based audio signal'.
  • a "multi-channel audio signal” may refer to an audio signal in which a plurality of audio signals are predefined as to which channel to output according to a channel layout. Furthermore, a “multi-channel audio signal” may also be referred to as a “channel-based audio signal” or a "channel audio signal.”
  • an "object audio signal” may refer to an audio signal that independently includes audio content and spatial information about the audio content, without predefining through which channel the audio signal will be output.
  • the "object audio signal” may be referred to as an "object-based audio signal.”
  • compression may refer to an operation of reducing the amount of data in an audio signal, thereby enabling more efficient storage or transmission of the audio signal.
  • compression may be performed by a codec or encoder, and the decompression operation may be performed by a codec or decoder.
  • 'mixing' means a signal processing operation of generating a new audio signal by adding the respective values obtained by multiplying each of a plurality of audio signals by their respective corresponding weights (i.e., mixing a plurality of audio signals).
  • FIG. 1 is a drawing schematically illustrating the operation of electronic devices according to one embodiment of the present disclosure.
  • an electronic device may perform encoding or decoding on an audio signal including an object audio signal.
  • an electronic device for encoding an audio signal and an electronic device for decoding an audio signal are each independent and separate electronic devices, and may be referred to as an encoding device (or encoder) and a decoding device (or decoder), respectively.
  • an electronic device for encoding an audio signal and an electronic device for decoding an audio signal may be a single electronic device (or codec), rather than independent and separate electronic devices.
  • encoding and decoding are described as being performed on a single object audio signal, but encoding and decoding may be performed on each of a plurality of object audio signals in the same manner.
  • the encoding device may perform encoding on an object audio signal (110) to obtain a bitstream (140).
  • the object audio signal (110) may include content audio.
  • Content audio refers to pure audio data corresponding to a dry signal excluding spatial information or metadata, and may be referred to by various expressions representing the same/similar concepts. For example, it may be replaced with expressions such as source audio, dry audio, and signal data, and is not limited to the examples described above.
  • the encoding device can obtain metadata (120) of the object audio signal (110) by rendering (e.g., object audio rendering) the object audio signal (110).
  • the encoding device can obtain metadata (120) including channels on which the object audio signal (110) is panned and gain values for each panned channel as a result of the rendering.
  • the encoding device can obtain metadata (120) including information that the channels to which a specific sample (e.g., a sample at time t1) of the object audio signal (110) is panned in a 7.1.4 channel layout (121) are a Lss (Left Surround Side) channel, a Lrs (Left Rear Surround) channel, and a Ltb (Left Top Back) channel, and that the gain values of the panned channels are 0.5, 0.3, and 0.1, respectively.
  • the encoding device can obtain metadata (120) including information about the channels to which a specific sample of the audio signal (110) is panned and the gain values of the panned channels, for each of a 5.1 channel layout (122) and a stereo channel layout (123) other than the 7.1.4 channel layout (121).
  • the encoding device may compress the object audio signal (110), metadata (120), and bed channel audio signal (130) to generate a bitstream (140).
  • the bed channel audio signal (130) may include a multi-channel audio signal of a specific layout that has been mixed in advance.
  • the bed channel audio signal (130) may include audio signals for each of a plurality of channels of a 5.1 channel layout, but is not necessarily limited to the example described above.
  • the encoding device may transmit the bitstream (140) to a decoding device.
  • the decoding device may decode the bitstream (140) to obtain an output audio signal output through a playback loudspeaker layout (150).
  • the playback loudspeaker layout may refer to a multi-channel layout of a speaker system that outputs the output audio signal obtained through decoding the bitstream (140).
  • the playback loudspeaker layout may be a 5.1.2 channel layout.
  • the decoding device can decompress the bitstream (140) to obtain an object audio signal (110), metadata (120), and a bed channel audio signal (130).
  • the decoding device can identify a target multi-channel layout that minimizes loss when rendering the object audio signal (110) among a plurality of multi-channel layouts of the metadata (120). For example, the decoding device can identify a 7.1.4 channel layout (121), which is a higher layout than a playback loudspeaker layout (150), among a 7.1.4 channel layout (121), a 5.1 channel layout (122), and a stereo channel layout (123) of the metadata (120), as the target multi-channel layout.
  • the decoding device can perform channel mixing to distribute the object audio signal (110) to a target multi-channel layout based on metadata (120).
  • the decoding device can render the multi-channel audio signal and the bed channel audio signal (130) obtained through the channel mixing to a playback loudspeaker layout (150).
  • the decoding device can obtain an output audio signal by mixing the rendered bed channel audio signal and the rendered multi-channel audio signal.
  • the output audio signal can be output through the playback loudspeaker layout (150).
  • the number of object audio signals that can be processed at once may be limited due to limitations in the computational capacity of the decoding device.
  • the number of object audio signals that the decoding device can process at once may be further limited. Accordingly, the spatial resolution of the output audio signal obtained through the decoding device may be reduced, and the user's sense of immersion in the output audio signal may be reduced.
  • An encoding device may transmit a bitstream including metadata including panned channels and gain values for each of the panned channels for each of a plurality of multi-channel layouts.
  • a decoding device may render an object audio signal based on metadata obtained from a bitstream received from an encoding device to obtain an output audio signal.
  • the decoding device can render object audio signals with a smaller computational load than when rendering based on spatial information. This may mean that the number of object audio signals that the decoding device can process simultaneously increases. That is, according to one embodiment of the present disclosure, the spatial resolution of the output audio signal obtained through the decoding device is improved, thereby further increasing the user's sense of immersion in the audio signal.
  • FIG. 2 is a flowchart illustrating a method for encoding an audio signal by an electronic device according to one embodiment of the present disclosure.
  • FIG. 2 operations of an electronic device encoding an audio signal are schematically described, and a detailed description of each operation will be described with reference to the drawings that follow.
  • the electronic device described in FIG. 2 may correspond to an electronic device, encoding device, or encoder that performs the encoding method of the present disclosure.
  • the electronic device can render a plurality of object audio signals to obtain metadata of the plurality of object audio signals.
  • the metadata may include, for each of the plurality of multi-channel layouts, the channels on which the plurality of object audio signals are panned and gain values for each of the panned channels.
  • rendering of an object audio signal may refer to an operation or function of an electronic device obtaining information about channels to which an object audio signal is panned in a specific multi-channel layout and gain values for each of the panned channels.
  • rendering of an object audio signal may be referred to as object audio rendering.
  • a panned channel may refer to a channel from which an object audio signal is output among a plurality of channels in a specific multi-channel layout.
  • the gain values for each of the panned channels may refer to values that determine the volume level of an object audio signal output from each of the object-panned channels.
  • the electronic device can obtain a plurality of audio samples by sampling a plurality of object audio signals according to a preset sampling rate (e.g., 48 kHz).
  • the electronic device can render a plurality of object audio signals for a specific multi-channel layout by obtaining panned channels and gain values for each of the panned channels of the plurality of audio samples in a specific multi-channel layout.
  • the processing of audio signals described in the present disclosure can be understood as processing of each of the plurality of audio samples obtained as a result of sampling.
  • metadata of a plurality of object audio signals may be individually acquired for each of the plurality of object audio signals and processed or mapped and managed in association with each of the plurality of object audio signals.
  • the metadata may further include an identifier (e.g., a unique object ID, etc.) for identifying which object audio signal among the plurality of object audio signals corresponds to.
  • the electronic device may associate and process or map and manage the plurality of object audio signals and their corresponding metadata based on the identifier of the metadata.
  • the metadata of the plurality of object audio signals may be metadata defined for each frame at a preset time interval (e.g., 10 ms).
  • the metadata may further include a time stamp for identifying which frame of the object audio signal it corresponds to.
  • the electronic device may process or map and manage the object audio signal and the metadata by temporally associating them based on the time stamp of the metadata.
  • the metadata for a specific frame may be determined as metadata for a plurality of audio samples at a time close to the specific frame.
  • the metadata may include metadata defined for each of the plurality of audio samples of the object audio signals.
  • the electronic device can render a plurality of object audio signals into a plurality of multi-channel layouts based on spatial information of a plurality of objects corresponding to the plurality of object audio signals.
  • the spatial information of an object may include information about the spatial characteristics of the object.
  • the spatial information may include information about the location of the object.
  • Information about the location of the object may be expressed as (x, y, z) coordinates in three-dimensional space based on the location of the listener, or may be expressed as azimuth, elevation, and radius.
  • the spatial information may be referred to by replacing it with various expressions that represent the same/similar concepts.
  • the spatial information may be replaced with expressions such as object metadata, positional metadata, spatial metadata, and dynamic metadata, but is not limited to the examples described above.
  • the spatial information of an object may include spatial information defined for each frame, which is a preset time interval.
  • the electronic device may obtain spatial information of audio samples corresponding to the time between consecutive frames based on the spatial information of each consecutive frame, if the preset time interval is longer than the time interval according to the sampling rate. For example, the electronic device may interpolate spatial information of audio samples corresponding to the time between consecutive frames based on the spatial information of each consecutive frame, or determine spatial information of a frame closest to the time corresponding to a specific audio sample as the spatial information of the corresponding audio sample.
  • the multiple multi-channel layouts may include a 7.1.4 channel layout, a 5.1 channel layout, and a stereo channel layout.
  • a multi-channel layout may refer to a configuration in which the spatial arrangement and roles of a plurality of speakers outputting audio signals are defined.
  • the multi-channel layout may be defined in an X.Y.Z format.
  • X may correspond to a number of channels corresponding to the number of speakers in a horizontal direction
  • Y may correspond to a number of low-frequency effects (LFE) channels
  • Z may correspond to a number of channels corresponding to the number of upper (height) speakers.
  • the positions of the plurality of channels of the multi-channel layout may be understood to correspond to positions of the plurality of speakers outputting audio signals
  • the plurality of channels may be arranged at positions defined according to a standard with respect to a listener.
  • a 7.1.4 channel layout may include seven horizontal channels, one low-frequency effects channel, and four top (height) channels.
  • the seven horizontal channels may include a Front Left (FL) channel located to the left in front of the listener, a Front Right (FR) channel located to the right in front of the listener, a Center (C) channel located centered in front of the listener, a Left Surround Side (Lss) channel located to the left side of the listener, a Right Surround Side (Rss) channel located to the right side of the listener, a Left Rear Surround (Lrs) channel located to the left rear of the listener, and a Right Rear Surround (Rrs) channel located to the right rear of the listener.
  • FL Front Left
  • FR Front Right
  • C Center
  • Lss Left Surround Side
  • Rss Right Surround Side
  • Lrs Left Rear Surround
  • Rrs Right Rear Surround
  • the four top channels of a 7.1.4 channel layout may include a Left Top Front (Ltf) channel located at the front left above the listener's head, a Right Top Back (Rtf) channel located at the front right above the listener's head, a Left Top Back (Ltb) channel located at the rear left above the listener's head, and a Right Top Back (Rtb) channel located at the rear right above the listener's head.
  • Ltf Left Top Front
  • Rtf Right Top Back
  • Ltb Left Top Back
  • Rtb Right Top Back
  • the low-frequency effects channels of a 7.1.4 channel layout may be positioned in any appropriate location regardless of the listener's position.
  • a 5.1 channel layout may include five horizontal channels and one low-frequency effects channel.
  • the five horizontal channels may include a Front Left (FL) channel located to the left in front of the listener, a Front Right (FR) channel located to the right in front of the listener, a Center (C) channel located to the center in front of the listener, a Left Side (Ls) channel located to the left side of the listener, and a Right Side (Rs) channel located to the right side of the listener.
  • the low-frequency effects channel of the 5.1 channel layout may be positioned in any suitable space regardless of its position relative to the listener.
  • a stereo channel layout may include two horizontal channels, such as a Left (L) channel located to the left in front of the listener and a Right (R) channel located to the right in front of the listener.
  • L Left
  • R Right
  • the multiple multi-channel layouts are not limited to the aforementioned 7.1.4 channel layout, 5.1 channel layout, and stereo channel layout. Some of these multi-channel layouts may be excluded, or other multi-channel layouts may be added. Furthermore, the names of the multiple channels of the multi-channel layout may differ from the examples described above.
  • the electronic device can identify channels to which the plurality of object audio signals are panned and gain values for each of the panned channels for each of the plurality of multi-channel layouts based on spatial information of the plurality of objects corresponding to the plurality of object audio signals.
  • the electronic device can identify the positions of speakers corresponding to multiple channels relative to a listener for a specific multi-channel layout.
  • the positions of the speakers can be expressed as (x, y, z) coordinates in three-dimensional space. For example, if the preset multi-channel layout is a stereo channel layout, the speaker corresponding to the L channel can be identified as being located at (-2, 0, 0) and the speaker corresponding to the R channel can be identified as being located at (2, 0, 0) relative to the listener.
  • the electronic device can identify the location of an object based on spatial information of the object. In one embodiment, the electronic device can compare the location of the object with the location of a speaker. In one embodiment, the electronic device can calculate the distance between the location of the object and the location of the speaker, and identify the speaker closest to the object or a speaker located within a preset distance based on the calculated distance. Here, the closest speakers can be identified in multiple numbers according to a multi-channel layout. In one embodiment, the electronic device can identify the channel corresponding to the identified speaker as the channel to which the object audio signal is panned.
  • the electronic device can identify a gain value of a panned channel based on a distance between a position of an object and a position of a speaker corresponding to the panned channel. In one embodiment, the electronic device can calculate a gain value of the panned channel such that the gain value decreases as the distance between the position of the object and the position of the speaker corresponding to the panned channel increases, and the gain value increases as the distance between the position of the object and the position of the speaker corresponding to the panned channel decreases. In one embodiment, the electronic device can normalize the gain values of the panned channels such that the gain values have values between 0 and 1.
  • the electronic device can identify the channels to which the object audio signal is panned and the gain values for each of the panned channels for each of the plurality of multi-channel layouts according to the above-described method, and store the identified panned channels and the gain values for each of the panned channels in the metadata of the object audio signal.
  • the electronic device can generate a bitstream by compressing a plurality of object audio signals, metadata, and bed channel audio signals.
  • the electronic device can compress a plurality of object audio signals, metadata of the plurality of object audio signals, and a bed channel audio signal into a binary format.
  • the electronic device can store the compressed plurality of object audio signals and their corresponding compressed metadata in a bitstream in a temporally synchronized state.
  • the electronic device can also store the compressed bed channel audio signal in the bitstream.
  • the electronic device can transmit the generated bitstream to an external electronic device.
  • the metadata may include mixing gain values for the plurality of object audio signals.
  • the mixing gain values for the plurality of object audio signals may be used by an electronic device that decodes an audio signal to mix a multi-channel audio signal obtained from the plurality of object audio signals.
  • the metadata may be metadata for each of a plurality of groups of a plurality of object audio signals.
  • the plurality of object audio signals may be grouped into a plurality of groups based on a user input.
  • the electronic device may obtain metadata for the plurality of object audio signals as metadata for each of the plurality of groups.
  • the metadata for each of the plurality of groups may include the channels to which each of the object audio signals included in the plurality of groups is panned and the gain values for each of the panned channels.
  • the metadata may further include a mixing gain value for each of the subgroups of the plurality of objects. The specific details of this will be described later through the operation of an electronic device that decodes audio signals.
  • FIG. 3 is a diagram for explaining an operation of an electronic device encoding an audio signal according to one embodiment of the present disclosure.
  • an electronic device (300) can perform encoding on an audio signal.
  • the electronic device (300) of FIG. 3 may correspond to the electronic device encoding an audio signal of FIG. 2.
  • the electronic device (300) may obtain a raw audio signal (310) on which encoding is performed.
  • the raw audio signal (310) may include a bed channel audio signal (312) and a plurality of object audio signals (314).
  • the raw audio signal (310) may be data stored in a memory of the electronic device or data received from an external electronic device.
  • the raw audio signal (310) may include an audio signal sampled according to a preset sampling rate.
  • the bed channel audio signal (312) and the plurality of object audio signals (314) may be composed of a plurality of audio samples sampled according to 48 kHz.
  • the electronic device (300) may obtain a user input (320) for determining a plurality of multi-channel layouts. For example, the electronic device (300) may determine the plurality of multi-channel layouts as a 7.1.4 channel layout, a 5.1 channel layout, and a stereo channel layout based on obtaining a user input (320) for selecting the plurality of multi-channel layouts as a 7.1.4 channel layout, a 5.1 channel layout, and a stereo channel layout. In one embodiment, the electronic device (300) may display a UI including a plurality of multi-channel layout candidates that the user can select. In one embodiment, the electronic device (300) may determine a multi-channel layout selected from among the plurality of multi-channel layout candidates included in the UI as the plurality of multi-channel layouts based on the user input.
  • the electronic device (300) may perform object audio rendering for a plurality of object audio signals (314) (S330). In one embodiment, the electronic device (300) may perform object audio rendering based on a plurality of object audio signals (314) and a plurality of multi-channel layouts. In one embodiment, the electronic device (300) may perform object audio rendering to obtain metadata including channels on which the plurality of object audio signals (314) are panned and gain values for each of the panned channels for each of the plurality of multi-channel layouts.
  • the electronic device (300) can generate a bitstream (340) by compressing object audio signals (314), metadata of the object audio signals (314), and bed channel audio signals (312). In one embodiment, the electronic device (300) can transmit the bitstream (340) to an external electronic device (e.g., an encoding device).
  • an external electronic device e.g., an encoding device
  • FIG. 4 is a diagram for explaining metadata of an object audio signal according to one embodiment of the present disclosure.
  • the electronic device (300) may perform object audio rendering on an object audio signal (410) to obtain metadata (420) of the object audio signal (410).
  • the metadata (420) may be processed or mapped and managed in association with the object audio signal (410).
  • the metadata (420) may be stored corresponding to each of a plurality of samples of the sampled object audio signal (410).
  • FIG. 4 describes that object audio rendering is performed on one object audio signal (410), object audio rendering may be performed in the same manner on each of a plurality of object audio signals, thereby obtaining metadata for each of the plurality of object audio signals.
  • the electronic device (300) can render the object audio signal (410) into a plurality of multi-channel layouts (414) based on spatial information (412) of an object corresponding to the object audio signal (410).
  • the spatial information (412) of the object may include information on the spatial characteristics of the object corresponding to the object audio signal (410).
  • 'Cartesian' may be information indicating the spatial location of the object in (x, y, z) coordinates.
  • 'screenRef' may be information indicating whether the location of the object is fixed to a specific location based on the screen.
  • 'ScreenEdgeLock' may be information indicating whether the object is fixed to the edge of the screen.
  • 'ChannelLock' may be information indicating whether to fix the channel through which the object audio signal is output.
  • 'ObjectDivergence' may be information indicating the degree of distribution when the object audio signal is simultaneously distributed and output to multiple speakers.
  • 'Width/Height/Depth(size)' may be information indicating the size of the object.
  • 'ZoneExclusion' may be information for setting the object audio signal not to be heard in a specific area (Zone).
  • 'Gain' may be information indicating the intensity at which an object audio signal is output.
  • 'Diffuse' may be information indicating how diffuse an object audio signal is heard.
  • the present invention is not necessarily limited to the above-described examples, and the object's spatial information (412) may further include various information for rendering the object audio signal.
  • the electronic device (300) can identify the channels to which the object audio signal (410) is panned and the gain values for each of the panned channels for each of the plurality of multi-channel layouts (414) based on the spatial information (412) of the object.
  • object audio rendering may be performed for a 7.1.4 channel layout among multiple multi-channel layouts (414).
  • the electronic device (300) may identify the positions of speakers corresponding to multiple channels of the 7.1.4 channel layout.
  • the electronic device (300) may identify the position of the object based on the 'Cartesian' of the spatial information (412) of the object.
  • the electronic device (300) may calculate the distance between the object and the speakers corresponding to multiple channels of the 7.1.4 channel layout. Based on the calculated distance, the electronic device (300) may identify the Lss channel, the Lrs channel, and the Ltb channel among the multiple channels of the 7.1.4 channel layout as channels corresponding to the three speakers closest to the object.
  • the electronic device (300) may identify the identified Lss channel, the Lrs channel, and the Ltb channel as channels to which the object audio signal (410) is panned in the 7.1.4 channel layout.
  • the electronic device (300) can store information about the identified Lss channel, Lrs channel, and Ltb channel in metadata (420).
  • the electronic device (300) can identify the gain values of the Lss channel, the Lrs channel, and the Ltb channel based on the distance between the position of the object and the position of the speaker corresponding to the Lss channel, the Lrs channel, and the Ltb channel.
  • the gain values of the Lss channel, the Lrs channel, and the Ltb channel can be calculated to be inversely proportional to the distance between the speaker and the position.
  • the gain value of the panned channel can be calculated according to the following formula.
  • G may correspond to a gain value
  • d may correspond to a distance between the position of the speaker and the position of the object.
  • the gain value of the Lss channel located closest to the object may be identified as 0.5
  • the gain value of the Lrs channel located second closest may be identified as 0.3
  • the gain value of the Ltb channel located third closest may be identified as 0.1.
  • the electronic device (300) may normalize the gain values of the Lss channel, the Lrs channel, and the Ltb channel so that the gain values have values between 0 and 1.
  • the electronic device (300) may store the gain values of the identified Lss channel, the Lrs channel, and the Ltb channel in the metadata (420).
  • the formula for calculating the gain value is not necessarily limited to the above-described example.
  • the electronic device (300) can adjust the gain values for each panned channel based on the 'ObjectDivergence' of the spatial information (412). For example, the gain value of the panned channel can be adjusted according to the following formula.
  • G' corresponds to the adjusted gain value
  • D corresponds to the degree of dispersion of the object
  • N corresponds to the number of panned channels.
  • the formula for adjusting the gain value is not necessarily limited to the example described above.
  • the electronic device (300) can identify the channels to which the object audio signal (410) is panned and the gain values for each panned channel for the 5.1 channel layout and the stereo channel layout, which are the remaining layouts except the 7.1.4 channel layout among the plurality of multi-channel layouts (414).
  • the present invention is not necessarily limited to the above-described example, and information about other spatial characteristics included in the spatial information (412) of the object other than the above-described spatial characteristics (the position and degree of dispersion of the object) may be further taken into consideration to identify the channels to which the object audio signal (410) is panned and the gain values for each panned channel.
  • the electronic device (300) may obtain metadata (420) of the object audio signal (410) based on a user input defining spatial characteristics of an object corresponding to the object audio signal (410). In one embodiment, the electronic device (300) may obtain user input defining spatial characteristics of an object corresponding to the object audio signal (410) to generate content including the object audio signal (410). For example, the electronic device (300) may obtain user input defining an object corresponding to a specific raw audio signal and determining a position or movement path of the defined object over time. In addition, the electronic device (300) may obtain user input determining a diffusion degree of the defined object. In this case, the electronic device (300) may obtain metadata (420) of the object audio signal (410) based on the spatial characteristics corresponding to the user input. In other words, the electronic device (300) can obtain metadata (420) of the object audio signal (410) based on user input defining spatial characteristics of the object rather than spatial information (412) of the object from the stage of generating content including the object audio signal (410).
  • the electronic device (300) may obtain metadata (420) of the object audio signal (410) based on user input including information about the channels to which the object audio signal (410) is panned and gain values for each of the panned channels. That is, when information included in the metadata (420) of the object audio signal (410) is directly input from a user, the electronic device (300) may obtain metadata (420) for the object audio signal (410) based on the user input without performing object audio rendering for the object audio signal (410).
  • metadata (420) may be used by an electronic device that performs decoding of an audio signal, which will be described later, to render an object audio signal (410).
  • the amount of computation required to render the object audio signal (410) may be reduced, as rendering of the object audio signal (410) based on spatial information (412) of the object does not need to be performed during the decoding step of the audio signal.
  • FIG. 5 is a diagram illustrating a method for an electronic device according to an embodiment of the present disclosure to decode an audio signal.
  • FIG. 5 operations of an electronic device decoding an audio signal are schematically described, and detailed descriptions of each operation will be described later with reference to the drawings.
  • the electronic device described in FIG. 5 may correspond to an electronic device, decoding device, or decoder that performs the decoding method of the present disclosure.
  • the electronic device can decompress the bitstream to obtain a plurality of object audio signals, metadata of the plurality of object audio signals, and a bed channel audio signal.
  • the bitstream may correspond to a bitstream received from an electronic device encoding an audio signal of FIG. 2.
  • the electronic device may receive a bitstream from an electronic device encoding an audio signal and decompress the received bitstream.
  • the metadata may include, for each of a plurality of multi-channel layouts, channels on which a plurality of object audio signals are panned and gain values for each of the panned channels.
  • the electronic device may interpolate metadata for each of a plurality of audio samples included between adjacent frames based on metadata of adjacent frames, when the metadata is defined in units of frames, which are preset time intervals. The metadata for each of the interpolated plurality of audio samples may be used in operations of the electronic device, which will be described later.
  • the electronic device can identify a target multi-channel layout that minimizes loss when rendering channel audio as a playback loudspeaker layout among multiple multi-channel layouts of metadata.
  • the playback loudspeaker layout may refer to a multi-channel layout of a speaker system that outputs an output audio signal obtained by decoding a bitstream.
  • the electronic device may obtain information about the playback loudspeaker layout from the speaker system that outputs the output audio signal.
  • the speaker system may be composed of a plurality of speakers that constitute the playback loudspeaker layout.
  • the electronic device may determine the playback loudspeaker layout based on a user input that includes information about the playback loudspeaker layout. Meanwhile, the speaker system may be referred to by various expressions that represent the same/similar concepts.
  • the speaker system may be replaced with expressions such as a multi-channel speaker configuration, a speaker arrangement, and an audio output system, and is not limited to the examples described above.
  • the speaker system may be a component included in the electronic device, or a separate component that performs wired or wireless communication with the electronic device.
  • channel audio rendering may refer to an operation of converting (or adjusting, mapping, etc.) an input multi-channel audio signal into a multi-channel audio signal of a specific multi-channel layout.
  • channel audio rendering may convert a multi-channel audio signal of a 7.1.4 channel layout into a multi-channel audio signal of a 7.1 channel layout.
  • channel audio rendering may refer to an audio signal of a specific multi-channel layout being spatially distributed to another multi-channel layout.
  • spatially distributing an audio signal may refer to mapping an audio signal to an output channel so as to correspond to a spatial direction of each channel based on spatial characteristics of the audio signal or positional information of a sound source.
  • the electronic device may identify a layout identical to the playback loudspeaker layout among a plurality of multi-channel layouts as the target multi-channel layout.
  • the layout identical to the playback loudspeaker layout may not cause any loss in performing channel audio rendering.
  • loss may refer to distortion or reduction of acoustic information that occurs in the process of converting a multi-channel audio signal with a higher number of channels into a multi-channel audio signal with a lower number of channels.
  • Loss may be referred to by various expressions that represent the same or similar concepts. For example, loss may be replaced with expressions such as 'audio loss', 'reduction of acoustic information', 'loss of spatial information', 'degradation of channel separation', 'degradation in audio quality', 'signal distortion', etc., and is not limited to the examples described above.
  • loss may also occur in the process of converting a multi-channel audio signal with a lower number of channels into a multi-channel audio signal with a higher number of channels.
  • the electronic device may identify a layout that includes a greater number of channels than the playback loudspeaker layout as the target multi-channel layout if there is no layout among the plurality of multi-channel layouts that is identical to the playback loudspeaker layout. In one embodiment, if there are multiple layouts among the plurality of multi-channel layouts that include a greater number of channels than the playback loudspeaker layout, the electronic device may identify a layout that includes a channel number closest to the channel number of the playback loudspeaker layout as the target multi-channel layout.
  • the layout that includes a greater number of channels than the number of channels of the playback loudspeaker layout while having the smallest difference in the number of channels may perform channel audio rendering through relatively less lossy down-mixing.
  • the electronic device may identify the layout with the largest number of channels among the plurality of multi-channel layouts as the target multi-channel layout.
  • the layout with the largest number of channels among the plurality of multi-channel layouts may be less likely to incur loss during channel audio rendering, as it has the largest amount of spatial acoustic information compared to other layouts.
  • the electronic device can perform channel mixing on a plurality of object audio signals based on metadata to obtain a first multi-channel audio signal of a target multi-channel layout.
  • an electronic device may obtain a plurality of object audio signals channelized into panned channels by applying a gain value to each panned channel for a target multi-channel layout to a plurality of object audio signals.
  • applying a gain value may mean multiplying the amplitude of an audio signal by the gain value to obtain an audio signal with an adjusted amplitude.
  • the electronic device can obtain a first multi-channel audio signal by combining a plurality of channelized object audio signals.
  • combining the audio signals may mean obtaining an audio signal in which the amplitudes of different audio signals are combined, and such combining may be performed on a sample-by-sample basis.
  • the electronic device can obtain a first object audio signal channelized into channels in which the first object audio signal is panned based on metadata of the first object audio signal, and can obtain a second object signal channelized into channels in which the second object audio signal is panned based on metadata of the second object audio signal.
  • the electronic device can obtain a first multi-channel audio signal of a target multi-channel layout by combining the channelized first object audio signal and the channelized second object audio signal.
  • the electronic device can perform channel audio rendering on the first multi-channel audio signal to obtain a second multi-channel audio signal of the playback loudspeaker layout.
  • the electronic device can obtain a second multi-channel audio signal by performing down-mixing on the first multi-channel audio signal if the target multi-channel layout has a larger number of channels than the playback loudspeaker layout. In one embodiment, the electronic device can obtain a second multi-channel audio signal by performing up-mixing on the first multi-channel audio signal if the target multi-channel layout has a smaller number of channels than the playback loudspeaker layout. In one embodiment, the electronic device can obtain the first multi-channel audio signal as the second multi-channel audio signal if the target multi-channel layout is the same as the playback loudspeaker layout.
  • the electronic device can perform channel audio rendering on a bed channel audio signal to obtain a third multi-channel audio signal of a playback loudspeaker layout.
  • the operation of the electronic device performing channel audio rendering of the multi-channel audio signal of the bed channel audio signal to the playback loudspeaker layout may correspond to the operation of performing channel audio rendering of the first multi-channel audio signal to the playback loudspeaker layout described above.
  • the electronic device may obtain an output audio signal by mixing a second multi-channel audio signal and a third multi-channel audio signal.
  • the output audio signal may include a multi-channel audio signal of a playback loudspeaker layout.
  • the output audio signal may be transmitted to a speaker system configured with a playback loudspeaker layout.
  • the electronic device can obtain an output audio signal by applying a mixing gain value of metadata to a second multi-channel audio signal.
  • the electronic device can apply the mixing gain value to the second multi-channel audio signal when mixing the second multi-channel audio signal and the third multi-channel audio signal, and obtain an output audio signal by combining the second multi-channel audio signal and the third multi-channel audio signal to which the mixing gain value is applied.
  • the metadata may further include a mixing gain value for each subgroup of the plurality of object audio signals.
  • the electronic device may obtain an output audio signal by applying the mixing gain value to the second multi-channel audio signal obtained by dividing it into each subgroup. This will be further described in detail in FIG. 8 below.
  • steps S530 and S540 may be performed for each of the subgroups of the plurality of object audio signals.
  • the electronic device may obtain a second multi-channel audio signal of the first subgroup by performing steps S530 and S540 on the object audio signals included in the first subgroup, and may obtain a second multi-channel audio signal of the second subgroup by performing steps S530 and S540 on the object audio signals included in the second subgroup.
  • the electronic device may obtain an output audio signal by applying a mixing gain value of metadata of the first subgroup to the second multi-channel audio signal of the first subgroup, and by applying a mixing gain value of metadata of the second subgroup to the multi-channel audio signal of the second subgroup.
  • FIG. 6 is a diagram for explaining an operation of an electronic device decoding an audio signal according to one embodiment of the present disclosure.
  • an electronic device (600) may perform decoding on an audio signal.
  • decoding on an audio signal may mean decoding a bitstream (610), which is an encoded audio signal.
  • the electronic device (600) of FIG. 6 may correspond to the electronic device for decoding an audio signal of FIG. 5.
  • the electronic device (600) may obtain a bitstream (610) on which decoding is performed.
  • the bitstream (610) may correspond to the bitstream (340) of FIG. 3.
  • the bitstream (610) may be data stored in a memory of the electronic device (600) or data received from an external electronic device.
  • the electronic device (600) can decompress the bitstream (610) to obtain a decoded audio signal (620).
  • the decoded audio signal (620) can include a bed channel audio signal (622) and a plurality of object audio signals (624).
  • the decoded audio signal (620) can include an audio signal sampled according to a preset sampling rate.
  • the bed channel audio signal (622) and the plurality of object audio signals (624) can be composed of a plurality of audio samples sampled according to 48 kHz.
  • the electronic device (600) may decompress the bitstream (610) to obtain metadata of a plurality of object audio signals (624).
  • the electronic device may associate each of the plurality of object audio signals (624) with its corresponding metadata and process or map and manage them.
  • the metadata may include, for each of the plurality of multi-channel layouts, channels on which the plurality of object audio signals (624) are panned and gain values for each of the panned channels.
  • the electronic device (600) may perform channel mixing on a plurality of object audio signals (624) (S630). In one embodiment, the electronic device (600) may perform channel mixing based on the plurality of object audio signals (624), metadata, and a playback loudspeaker layout to obtain a first multi-channel audio signal of a target multi-channel layout.
  • the first multi-channel audio signal may include an audio signal in which the plurality of object audio signals (624) are channelized into panned channels of the target multi-channel layout.
  • the first multi-channel audio signal may include a multi-channel audio signal in which the plurality of object signals (624) are channelized into the target multi-channel layout and combined for each channel of the target multi-channel layout.
  • the electronic device (600) can identify a target multi-channel layout among multiple multi-channel layouts of metadata that minimizes loss when rendering channel audio as a playback loudspeaker layout. In one embodiment, the electronic device (600) can obtain information about the playback loudspeaker layout from a speaker system (650) that outputs an output audio signal. In one embodiment, the playback loudspeaker layout can correspond to the layout of the speaker system (650).
  • the electronic device (600) may perform channel audio rendering on a first multi-channel audio signal (S642). In one embodiment, the electronic device (600) may perform channel audio rendering based on the first multi-channel audio signal and a playback loudspeaker layout to obtain a second multi-channel audio signal of the playback loudspeaker layout.
  • the second multi-channel audio signal may be a multi-channel audio signal in which a plurality of object audio signals (624) are converted to fit the playback loudspeaker layout and spatially distributed to each channel.
  • the electronic device (600) may perform channel audio rendering on the bed channel audio signal (622) (S644). In one embodiment, the electronic device (600) may perform channel audio rendering based on the bed channel audio signal and the playback loudspeaker layout to obtain a third multi-channel audio signal of the playback loudspeaker layout.
  • the third multi-channel audio signal may be a multi-channel audio signal in which the bed channel audio signal (622) is converted to fit the playback loudspeaker layout and spatially distributed to each channel.
  • the electronic device (600) can mix a second multi-channel audio signal and a third multi-channel audio signal (S650). In one embodiment, the electronic device (600) can obtain an output audio signal by mixing the second multi-channel audio signal and the third multi-channel audio signal. In one embodiment, the electronic device (600) can transmit the output audio signal to a speaker system (660). In one embodiment, the speaker system (660) can output the received output audio signal.
  • the output audio signal can be a multi-channel audio signal in which a bed channel audio signal (622) and a plurality of object audio signals (624) are converted to fit a playback loudspeaker layout and spatially distributed to each channel.
  • FIG. 7 is a diagram illustrating channel mixing and channel audio rendering according to one embodiment of the present disclosure.
  • the electronic device (600) can perform channel mixing based on an object audio signal (710), metadata (720), and a playback loudspeaker layout (730).
  • the object audio signal (710) and metadata (720) can correspond to one of the plurality of object audio signals (624) of FIG. 6 and metadata therefor.
  • the operations of the electronic device (600) described in FIG. 7 can be understood as operations for each of the plurality of object audio signals (624).
  • the electronic device (600) can identify a target multi-channel layout that minimizes loss when rendering channel audio as a playback loudspeaker layout (730) among the plurality of multi-channel layouts of the metadata (720). In one embodiment, the electronic device (600) can identify whether there is a layout identical to the playback loudspeaker layout (730) among the plurality of multi-channel layouts of the metadata (720). In one embodiment, if there is a layout identical to the playback loudspeaker layout (730) among the plurality of multi-channel layouts of the metadata (720), the electronic device (600) can identify the identical layout as the target multi-channel layout. For example, if there is a 7.1 channel layout that is the playback loudspeaker layout (730) among the plurality of multi-channel layouts of the metadata (720), the electronic device (600) can identify the 7.1 channel layout as the target multi-channel layout.
  • the electronic device (600) may identify a layout having a greater number of channels than the playback loudspeaker layout (730) as the target multi-channel layout if there is no layout identical to the playback loudspeaker layout (730) among the plurality of multi-channel layouts of the metadata (720). In one embodiment, the electronic device (600) may identify a layout having a largest number of channels among the plurality of multi-channel layouts as the target multi-channel layout if there is no layout identical to the playback loudspeaker layout (730) among the plurality of multi-channel layouts of the metadata (720).
  • the electronic device (600) may identify a 7.1.4 channel layout (721) that includes a greater number of channels than the 7.1 channel layout as the target multi-channel layout, or may identify a 7.1.4 channel layout (721) that includes the largest number of channels among the multiple multi-channel layouts of the metadata (720) as the target multi-channel layout.
  • the electronic device (600) can perform channel mixing on the object audio signal (710) to obtain a first multi-channel audio signal (740) of a target multi-channel layout. In one embodiment, the electronic device (600) can apply panned channels and gain values for each channel for the target multi-channel layout to the object audio signal (710) to obtain a channelized object audio signal with panned channels. In one embodiment, the electronic device (600) can combine the channelized object audio signals with panned channels to obtain the first multi-channel audio signal (740).
  • the electronic device (600) can identify, based on the metadata (720), that the panned channels for the 7.1 channel layout, which is the target multi-channel layout, are the Lss channel, the Lrs channel, and the Ltb channel, and that the gain values for each channel are 0.5, 0.3, and 0.1, respectively. Then, the electronic device (600) can obtain an audio signal obtained by applying the gain value 0.5 of the Lss channel to the object audio signal (710) as an object audio signal (711) channelized into the Lss channel. In the same manner, the electronic device (600) can obtain an object audio signal (712) channelized into the Lrs channel and an object audio signal (713) channelized into the Ltb channel.
  • the electronic device (600) can obtain channelized object signals for other object audio signals in the same manner.
  • the electronic device (600) can obtain a first multi-channel audio signal (740) by combining (or combining) multiple channelized object audio signals by channel.
  • the electronic device (600) can perform channel mixing on a first multi-channel audio signal (740) to obtain a second multi-channel audio signal (750) of a playback loudspeaker layout (730).
  • the electronic device (600) can perform downmixing to convert a first multi-channel audio signal (740) of a 7.1.4 channel layout into a second multi-channel audio signal (750) of a 7.1. channel layout.
  • the electronic device (600) can perform downmixing by distributing audio signals of four upper channels (Ltf, Rtf, Ltb, Rtb) of the 7.1.4 channel layout to seven horizontal channels (L, R, C, LFE, Lss, Rss, Lrs, Rrs).
  • the electronic device (600) can perform downmixing by applying a predefined panning algorithm or mixing matrix including gain values applied to each channel to the audio signals of the four upper channels (Ltf, Rtf, Ltb, Rtb).
  • the present invention is not necessarily limited to the above-described example, and if the target multi-channel layout (or the multi-channel layout of the first multi-channel audio signal (740)) is the same as the playback loudspeaker layout (730), channel audio rendering may not be performed. In addition, if the number of channels of the target multi-channel layout (or the number of channels of the first multi-channel audio signal (740)) is less than the number of channels of the playback loudspeaker layout (730), upmixing may be performed to convert the first multi-channel audio signal (740) into a second multi-channel audio signal (750).
  • FIG. 8 is a diagram for explaining an operation of an electronic device according to an embodiment of the present disclosure performing rendering of grouped object audio signals.
  • the electronic device (600) can perform rendering on a plurality of object audio signals grouped into a plurality of groups.
  • the metadata of the plurality of object audio signals may be metadata for each of the plurality of groups of the plurality of object audio signals.
  • the electronic device (600) can perform channel audio rendering of the object audio signals included in the plurality of groups based on the metadata of each of the plurality of groups.
  • the metadata for each of the plurality of groups may include panned channels and mixing gain values of a specific channel layout.
  • the electronic device (600) may obtain a first multi-channel audio signal of each of the plurality of groups by distributing object audio signals included in each of the plurality of groups to the panned channels of the metadata.
  • a plurality of object audio signals may be grouped into a first group (810-1) including a first object audio signal and a second object audio signal, a second group (810-2) including a third object audio signal, a fourth object audio signal, and a fifth object audio signal, and a third group (810-3) including a sixth object audio signal and a seventh object audio signal (810-3).
  • metadata of the plurality of object audio signals may include metadata (820-1) for the first group (810-1), metadata (820-2) for the second group (810-2), and metadata (820-3) for the third group (810-3).
  • metadata of each of the plurality of groups may include channels and mixing gain values to which the object audio signals included in the plurality of groups are panned.
  • the electronic device (600) can obtain the first multi-channel audio signal of the first group (810-1) by distributing the first object audio signal and the second object audio signal included in the first group (810-1) to the LTF channel and the RTF channel of the 7.1.4 channel layout of the metadata (820-1) of the first group (810-1).
  • the electronic device (600) can obtain the first multi-channel audio signal of the second group (810-2) and the first multi-channel audio signal of the third group (810-3) in the same manner.
  • the electronic device (600) can perform channel audio rendering on a first multi-channel audio signal of each of the plurality of groups to a playback loudspeaker layout to obtain a second multi-channel audio signal of each of the plurality of groups of playback loudspeaker layouts.
  • the electronic device (600) can perform channel audio rendering on the first multi-channel audio signals of the first to third groups (810-1 to 810-3) to a playback loudspeaker layout.
  • the electronic device (600) can perform upmixing or downmixing on the first multi-channel audio signals of the first to third groups (810-1 to 810-3) to a playback loudspeaker layout to obtain the second multi-channel audio signals of the first to third groups (810-1 to 810-3) from the first multi-channel audio signals.
  • the electronic device (600) can perform upmixing or downmixing on the bed channel audio signals to a playback loudspeaker layout to obtain the third multi-channel audio signals of the bed channel audio signals.
  • the electronic device (600) can obtain an output audio signal by applying a mixing gain value of metadata to a second multi-channel audio signal obtained by being divided into each of a plurality of groups.
  • the electronic device (600) can apply a mixing gain value of 0.85 of the metadata (820-1) of the first group (810-1) to the second multi-channel audio signal of the first group (810-1).
  • the electronic device (600) can apply a mixing gain value of 0.707 of the metadata (820-2) of the second group (810-2) to the second multi-channel audio signal of the second group (810-2).
  • the electronic device (600) can apply a mixing gain value of 1.0 of the metadata (820-3) of the third group (810-3) to the second multi-channel audio signal of the third group (810-3).
  • the electronic device (600) can obtain an output audio signal by combining the second multi-channel audio signal of the first to third groups (810-1 to 810-3) to which the mixing gain value is applied and the third multi-channel audio signal of the bed channel audio signal.
  • the metadata in FIG. 8 is illustrated as including only panned channels for one channel layout and not including gain values for each panned channel
  • the metadata may further include panned channels and gain values for each panned channel for each of a plurality of multi-channel layouts.
  • the electronic device (600) may further include a step of performing channel mixing for each of a plurality of groups of a plurality of object audio signals, as illustrated in FIG. 6.
  • the first multi-channel audio signals of the first group (810-1) to the third group (810-3) may be understood as multi-channel audio signals obtained by performing channel mixing of FIG. 6.
  • the electronic device (600) renders multiple object audio signals in groups and can individually apply a mixing gain value for each group.
  • the metadata since the metadata only needs to include panned channels and gain values for a single multi-channel layout, the size of the metadata required for rendering can be reduced.
  • the mixing gain value can be applied to each group in the final mixing stage, customized audio rendering can be performed according to importance or user preference.
  • FIG. 9 is a schematic diagram of an electronic device for encoding an audio signal according to one embodiment of the present disclosure.
  • the electronic device (900) may include a memory (910) and a processor (920).
  • the components illustrated in FIG. 9 are merely in accordance with one embodiment of the present disclosure, and the components included in the electronic device (900) are not limited to those illustrated in FIG. 9.
  • the electronic device (900) according to one embodiment of the present disclosure may not include some of the components illustrated in FIG. 9, and may further include components not illustrated in FIG. 9.
  • the electronic device (900) of FIG. 9 may correspond to the electronic device (300) of FIG. 3. In other words, the electronic device (900) of FIG. 9 may correspond to an electronic device that performs encoding on an audio signal.
  • the memory (910) may store instructions or program codes for performing functions or operations of the electronic device (2000).
  • at least one instruction, algorithm, data structure, program code, and application program stored in the memory (910) may be implemented in a programming or scripting language such as, for example, C, C++, Java, or an assembler.
  • the memory (910) may include at least one of a flash memory type, a hard disk type, a multimedia card micro type, a card type memory (e.g., SD or XD memory, etc.), a random access memory (RAM), a static random access memory (SRAM), a read-only memory (ROM), an electrically erasable programmable read-only memory (EEPROM), a programmable read-only memory (PROM), a mask ROM, a flash ROM, a hard disk drive (HDD), or a solid state drive (SSD).
  • the memory (910) may not exist separately and may be configured to be included in the processor (920).
  • the memory (910) may be configured as a volatile memory, a nonvolatile memory, or a combination of a volatile memory and a nonvolatile memory.
  • a program or at least one instruction for performing operations according to embodiments described below may be stored in the memory (910).
  • the memory (910) may also provide stored data to the processor (920) at the request of the processor (920).
  • the memory (910) may include an object audio rendering module (911).
  • the modules stored in the memory (910) refer to a unit that processes one function or operation of the electronic device (900), which may be implemented as hardware included in the electronic device (900) or software stored in the electronic device (900) or a combination of hardware and software. That is, the operation and function of the modules stored in the memory (910) may be understood as the operation and function of the electronic device (900).
  • the object audio rendering module (911) can perform object audio rendering for a plurality of object audio signals. In one embodiment, the object audio rendering module (911) can perform object audio rendering based on a plurality of object audio signals and a plurality of multi-channel layouts to obtain metadata of the plurality of object audio signals. Since the operation and function of the object audio rendering module (911) can correspond to the operation and function of the electronic device (300) of FIG. 3 that performs object audio rendering, a redundant description will be omitted.
  • the processor (920) can control the overall operations of the electronic device (900).
  • the processor (920) can include multiple processors.
  • at least one processor (920) can perform the operations and functions of the electronic device (900) disclosed herein by executing one or more instructions of a program stored in the memory (910).
  • the processor (920) may be configured as at least one of, for example, a Central Processing Unit, a microprocessor, a Graphic Processing Unit, Application Specific Integrated Circuits (ASICs), Digital Signal Processors (DSPs), Digital Signal Processing Devices (DSPDs), Programmable Logic Devices (PLDs), Field Programmable Gate Arrays (FPGAs), an Application Processor, a Neural Processing Unit, or an artificial intelligence processor designed with a hardware structure specialized for processing artificial intelligence models, but is not limited thereto.
  • ASICs Application Specific Integrated Circuits
  • DSPs Digital Signal Processors
  • DSPDs Digital Signal Processing Devices
  • PLDs Programmable Logic Devices
  • FPGAs Field Programmable Gate Arrays
  • an Application Processor a Neural Processing Unit
  • an artificial intelligence processor designed with a hardware structure specialized for processing artificial intelligence models, but is not limited thereto.
  • the multiple operations may be performed by a single processor or by multiple processors.
  • the first operation, the second operation, and the third operation may all be performed by a first processor, or the first and second operations may be performed by a first processor and the third operation may be performed by a second processor.
  • the embodiments of the present disclosure are not limited thereto.
  • One or more processors according to the present disclosure may be implemented as a single-core processor or a multi-core processor. If a method according to an embodiment of the present disclosure includes multiple operations, the multiple operations may be performed by a single core or by multiple cores included in one or more processors.
  • At least one processor (920) can render a plurality of object audio signals by executing one or more instructions to obtain metadata of the plurality of object audio signals.
  • At least one processor (920) may execute one or more instructions to compress a plurality of object audio signals, metadata, and bed channel audio signals to generate a bitstream.
  • the metadata may include, for each of the plurality of multi-channel layouts, the channels on which the plurality of object audio signals are panned and gain values for the panned channels.
  • At least one processor can render a plurality of object audio signals into a plurality of multi-channel layouts based on spatial information of a plurality of objects corresponding to the plurality of object audio signals by executing one or more instructions.
  • the multiple multi-channel layouts may include a 7.1.4 channel layout, a 5.1 channel layout, and a stereo channel layout.
  • At least one processor can perform various operations and functions of the electronic device disclosed in the present specification by executing one or more instructions.
  • FIG. 10 is a schematic diagram of an electronic device for decoding an audio signal according to one embodiment of the present disclosure.
  • the electronic device (1000) may include a memory (1100) and a processor (1200).
  • the components illustrated in FIG. 10 are merely in accordance with one embodiment of the present disclosure, and the components included in the electronic device (1000) are not limited to those illustrated in FIG. 10.
  • the electronic device (1000) according to one embodiment of the present disclosure may not include some of the components illustrated in FIG. 10, and may further include components not illustrated in FIG. 10.
  • the electronic device (1000) of FIG. 10 may correspond to the electronic device (600) of FIG. 6.
  • the electronic device (1000) of FIG. 10 may correspond to an electronic device that performs decoding on an audio signal.
  • the memory (1100) may store commands or program codes for performing functions or operations of the electronic device (1000).
  • the memory (1100) of FIG. 10 may correspond to the memory (910) of FIG. 9, and any overlapping descriptions of the configuration, operation, and function will be omitted.
  • the memory (1100) may include at least one of a channel mixing module (1110), a channel audio rendering module (1120), and an output mixing module (1130).
  • the channel mixing module (1100) can perform channel mixing on a plurality of object audio signals. In one embodiment, the channel mixing module (1100) can perform channel mixing on a plurality of object audio signals based on metadata of the object audio signals to obtain a first multi-channel audio signal of a target multi-channel layout. Since the operation and function of the channel mixing module (1100) can correspond to the operation and function of the electronic device (600) of FIG. 6 that performs channel mixing, a redundant description thereof will be omitted.
  • the channel audio rendering module (1120) can perform channel audio rendering on a first multi-channel audio signal. In one embodiment, the channel audio rendering module (1120) can perform channel audio rendering on the first multi-channel audio signal to obtain a second multi-channel audio signal of a playback loudspeaker layout. In one embodiment, the channel audio rendering module (1120) can perform channel audio rendering on a bed channel audio signal. In one embodiment, the channel audio rendering module (1120) can perform channel audio rendering on the bed channel audio signal to obtain a third multi-channel audio signal of a playback loudspeaker layout. Since the operation and function of the channel audio rendering module (1120) can correspond to the operation and function of the electronic device (600) of FIG. 6 that performs channel audio rendering, a redundant description thereof will be omitted.
  • the output mixing module (1130) can perform mixing of the second multi-channel audio signal and the third multi-channel audio. In one embodiment, the output mixing module (1130) can obtain an output audio signal by mixing the second multi-channel audio signal and the third multi-channel audio signal. Since the operation and function of the output mixing module (1130) correspond to the operation and function of the electronic device (600) that mixes the second multi-channel audio signal and the third multi-channel audio signal, a redundant description will be omitted.
  • the processor (1200) can control the overall operations of the electronic device (1000).
  • the processor (1200) can include multiple processors.
  • at least one processor (1200) can perform the operations and functions of the electronic device (1000) disclosed in this specification by executing one or more instructions of a program stored in the memory (1100).
  • the processor (1200) of FIG. 10 may correspond to the processor (920) of FIG. 9, and any overlapping descriptions among the configuration, operation, and function will be omitted.
  • At least one processor (1200) can decompress a bitstream by executing one or more instructions to obtain a plurality of object audio signals, metadata of the plurality of object audio signals, and a bed channel audio signal.
  • At least one processor (1200) can identify a target multi-channel layout that minimizes loss when rendering channel audio as a playback loudspeaker layout among a plurality of multi-channel layouts of metadata by executing one or more instructions.
  • At least one processor (1200) may perform channel mixing on a plurality of object audio signals based on metadata by executing one or more instructions to obtain a first multi-channel audio signal of a target multi-channel layout.
  • At least one processor (1200) may perform channel audio rendering on a first multi-channel audio signal by executing one or more instructions to obtain a second multi-channel audio signal of a playback loudspeaker layout.
  • the metadata may include, for each of the plurality of multi-channel layouts, the channels on which the plurality of object audio signals are panned and gain values for the panned channels.
  • At least one processor (1200) can identify a layout among a plurality of multi-channel layouts that is identical to a playback loudspeaker layout as a target multi-channel layout by executing one or more instructions.
  • At least one processor (1200) can identify a layout including a greater number of channels than the playback loudspeaker layout as a target multi-channel layout by executing one or more instructions, if no layout among the plurality of multi-channel layouts is identical to the playback loudspeaker layout.
  • At least one processor (1200) can identify a layout having the largest number of channels among the plurality of multi-channel layouts as a target multi-channel layout by executing one or more instructions, if there is no layout identical to the playback loudspeaker layout among the plurality of multi-channel layouts.
  • the metadata may further include mixing gain values for each of the plurality of groups of the plurality of object audio signals grouped into the plurality of groups.
  • At least one processor (1200) may obtain an output audio signal by applying a mixing gain value to a second multi-channel audio signal obtained by being divided into each of a plurality of groups by executing one or more instructions.
  • FIG. 11 is a detailed configuration diagram of an electronic device according to one embodiment of the present disclosure.
  • the electronic device (2000) of FIG. 11 may constitute all or part of the electronic device (900) of FIG. 9, and may constitute all or part of the electronic device (1000) of FIG. 10.
  • the electronic device (2000) may include a memory (2100), a communication interface (2200), a display (2300), an input interface (2400), an output interface (2500), and a processor (2600).
  • the components illustrated in FIG. 11 are merely according to one embodiment of the present disclosure, and the components included in the electronic device (2000) are not limited to those illustrated in FIG. 11.
  • the electronic device (2000) according to one embodiment of the present disclosure may not include some of the components illustrated in FIG. 11, and may further include components not illustrated in FIG. 11.
  • the memory (2100) may store commands or program codes for performing functions or operations of the electronic device (2000).
  • the memory (2100) of FIG. 11 may correspond to the memory (910) of FIG. 9 or the memory (1100) of FIG. 10, and any overlapping descriptions of the configuration, operation, and function will be omitted.
  • the communication interface (2200) is a component for the electronic device (2000) to communicate with an external electronic device.
  • the communication interface (2300) may perform data communication between the electronic device (2000) and the external electronic device using at least one of data communication methods including wired LAN, wireless LAN, Wi-Fi, Bluetooth, zigbee, Wi-Fi Direct (WFD), infrared Data Association (IrDA), Bluetooth Low Energy (BLE), Near Field Communication (NFC), Wireless Broadband Internet (Wibro), World Interoperability for Microwave Access (WiMAX), Shared Wireless Access Protocol (SWAP), Wireless Gigabit Alliance (WiGig), and RF communication.
  • data communication methods including wired LAN, wireless LAN, Wi-Fi, Bluetooth, zigbee, Wi-Fi Direct (WFD), infrared Data Association (IrDA), Bluetooth Low Energy (BLE), Near Field Communication (NFC), Wireless Broadband Internet (Wibro), World Interoperability for Microwave Access (WiMAX), Shared Wireless Access Protocol (SWAP), Wireless Giga
  • the communication interface (2200) can receive at least one of an object audio signal, a bed channel audio signal, and spatial information of the object audio signal from an external electronic device. In one embodiment, the communication interface (2200) can transmit a bitstream to the external electronic device. In one embodiment, the communication interface (2200) can receive a bitstream from the external electronic device. In one embodiment, the communication interface (2200) can receive information about a playback loudspeaker layout from the external electronic device. In one embodiment, the communication interface (2200) can transmit an output audio signal to the external electronic device.
  • the present invention is not limited to the examples described above, and the communication interface (2200) can receive or transmit various data necessary for performing operations and functions of the electronic device disclosed herein from or to the external electronic device.
  • the display (2300) is a component for displaying images and/or videos.
  • the display (2300) may be configured as a physical device including at least one of a liquid crystal display, a thin film transistor-liquid crystal display, an organic light-emitting diode (OLED), a flexible display, a 3D display, and an electrophoretic display.
  • the display (2300) may display a UI for encoding or decoding an audio signal based on a signal received from the processor (2600).
  • the display (2300) may display a UI for obtaining a user input disclosed herein based on a signal received from the processor (2600).
  • the present invention is not limited to the above-described examples, and the display (2300) may display various graphic elements necessary to perform operations and functions of the electronic device disclosed herein based on signals received from the processor (2600).
  • the input interface (2400) is a component for receiving various user inputs.
  • the input interface (2400) may include a touch panel, a physical button, a microphone, etc.
  • information input through the input interface (2400) may be provided to the processor (2600).
  • the input interface (2400) may obtain a user input for grouping a plurality of object audio signals into a plurality of groups.
  • the input interface (2400) may obtain a user input for determining a plurality of multi-channel layouts.
  • the input interface (2400) may obtain a user input for defining spatial characteristics of an object corresponding to an object audio signal.
  • the input interface (2400) may obtain a user input including information about channels to which the object audio signal is panned and a gain value for each panned channel. In one embodiment, the input interface (2400) may obtain a user input including information about a playback loudspeaker layout. However, it is not necessarily limited to the above-described examples, and the input interface (2400) can obtain various data necessary to perform the operations and functions of the electronic device (2000) disclosed in this specification.
  • the output interface (2500) is a component for the electronic device (2000) to provide various information to the user.
  • the electronic device (2000) may include a plurality of speakers that are configured to output sounds.
  • the plurality of speakers may form a speaker system with a playback loudspeaker layout.
  • the plurality of speakers may output output audio signals.
  • the present invention is not necessarily limited to the above-described example, and the output interface (2500) may output various sounds or information for performing the operations and functions of the electronic device (2000) disclosed herein.
  • the processor (2600) can control the overall operations of the electronic device (2000).
  • the processor (2600) can include multiple processors.
  • at least one processor (2600) can perform the operations and functions of the electronic device (2000) disclosed in the present specification by executing one or more instructions of a program stored in the memory (2100).
  • the processor (2600) of FIG. 10 may correspond to the processor (920) of FIG. 9 or the processor (1200) of FIG. 10, and any overlapping descriptions among the configuration, operation, and function will be omitted.
  • Computer-readable media may be any available media that can be accessed by a computer, and include both volatile and nonvolatile media, removable and non-removable media.
  • Computer-readable media may include computer storage media and communication media.
  • Computer storage media include both volatile and nonvolatile, removable and non-removable media implemented in any method or technology for storing information, such as computer-readable instructions, data structures, program modules, or other data.
  • Communication media may typically include computer-readable instructions, data structures, or other data in a modulated data signal, such as program modules.
  • a computer-readable storage medium may be provided in the form of a non-transitory storage medium.
  • the term “non-transitory storage medium” simply means a tangible device that does not contain signals (e.g., electromagnetic waves). This term does not distinguish between cases where data is permanently stored in the storage medium and cases where data is temporarily stored.
  • a “non-transitory storage medium” may include a buffer in which data is temporarily stored.
  • the method according to various embodiments disclosed in the present document may be provided as included in a computer program product.
  • the computer program product may be traded as a product between a seller and a buyer.
  • the computer program product may be distributed in the form of a machine-readable storage medium (e.g., compact disc read-only memory (CD-ROM)), or may be distributed online (e.g., downloaded or uploaded) through an application store or directly between two user devices (e.g., smartphones).
  • a portion of the computer program product e.g., a downloadable app
  • a machine-readable storage medium such as the memory of a manufacturer's server, an application store's server, or an intermediary server.

Landscapes

  • Engineering & Computer Science (AREA)
  • Physics & Mathematics (AREA)
  • Multimedia (AREA)
  • Signal Processing (AREA)
  • Acoustics & Sound (AREA)
  • Audiology, Speech & Language Pathology (AREA)
  • Human Computer Interaction (AREA)
  • Health & Medical Sciences (AREA)
  • Computational Linguistics (AREA)
  • Theoretical Computer Science (AREA)
  • Mathematical Physics (AREA)
  • Library & Information Science (AREA)
  • Data Mining & Analysis (AREA)
  • Databases & Information Systems (AREA)
  • General Engineering & Computer Science (AREA)
  • General Physics & Mathematics (AREA)
  • Stereophonic System (AREA)

Abstract

L'invention concerne un dispositif électronique et un procédé de traitement d'un signal audio. Le dispositif électronique et le procédé de traitement d'un signal audio peuvent comprendre une étape de génération d'un flux binaire par compression d'une pluralité de signaux audio d'objet, de métadonnées des signaux audio d'objet et d'un signal audio de canaux de lit. Le dispositif électronique et le procédé de traitement d'un signal audio peuvent comprendre une étape d'obtention de la pluralité de signaux audio d'objet, des métadonnées de la pluralité de signaux audio d'objet, et du signal audio de canaux de lit par décompression du flux binaire. Les métadonnées peuvent comprendre, pour chaque configuration d'une pluralité de configurations multicanal, des canaux dans lesquels la pluralité de signaux audio d'objet sont panoramiques et une valeur de gain pour chacun des canaux dans lesquels la pluralité de signaux audio d'objet sont panoramiques.
PCT/KR2025/095221 2024-05-07 2025-04-16 Dispositif électronique et procédé de traitement de signal audio Pending WO2025234858A1 (fr)

Applications Claiming Priority (4)

Application Number Priority Date Filing Date Title
KR20240060111 2024-05-07
KR10-2024-0060111 2024-05-07
KR10-2024-0194621 2024-12-23
KR1020240194621A KR20250160809A (ko) 2024-05-07 2024-12-23 오디오 신호를 처리하는 전자 장치 및 방법

Publications (1)

Publication Number Publication Date
WO2025234858A1 true WO2025234858A1 (fr) 2025-11-13

Family

ID=97674966

Family Applications (1)

Application Number Title Priority Date Filing Date
PCT/KR2025/095221 Pending WO2025234858A1 (fr) 2024-05-07 2025-04-16 Dispositif électronique et procédé de traitement de signal audio

Country Status (1)

Country Link
WO (1) WO2025234858A1 (fr)

Citations (5)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
KR101681529B1 (ko) * 2013-07-31 2016-12-01 돌비 레버러토리즈 라이쎈싱 코오포레이션 공간적으로 분산된 또는 큰 오디오 오브젝트들의 프로세싱
KR102243395B1 (ko) * 2013-09-05 2021-04-22 한국전자통신연구원 오디오 부호화 장치 및 방법, 오디오 복호화 장치 및 방법, 오디오 재생 장치
KR102320279B1 (ko) * 2017-05-03 2021-11-03 프라운호퍼 게젤샤프트 쭈르 푀르데룽 데어 안겐반텐 포르슝 에. 베. 오디오 렌더링을 위한 오디오 프로세서, 시스템, 방법 및 컴퓨터 프로그램
CN116723438A (zh) * 2023-05-26 2023-09-08 三星电子(中国)研发中心 修正参数生成方法和装置
JP2024029123A (ja) * 2013-09-12 2024-03-05 ドルビー ラボラトリーズ ライセンシング コーポレイション ダウンミックスされたオーディオ・コンテンツについてのラウドネス調整

Patent Citations (5)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
KR101681529B1 (ko) * 2013-07-31 2016-12-01 돌비 레버러토리즈 라이쎈싱 코오포레이션 공간적으로 분산된 또는 큰 오디오 오브젝트들의 프로세싱
KR102243395B1 (ko) * 2013-09-05 2021-04-22 한국전자통신연구원 오디오 부호화 장치 및 방법, 오디오 복호화 장치 및 방법, 오디오 재생 장치
JP2024029123A (ja) * 2013-09-12 2024-03-05 ドルビー ラボラトリーズ ライセンシング コーポレイション ダウンミックスされたオーディオ・コンテンツについてのラウドネス調整
KR102320279B1 (ko) * 2017-05-03 2021-11-03 프라운호퍼 게젤샤프트 쭈르 푀르데룽 데어 안겐반텐 포르슝 에. 베. 오디오 렌더링을 위한 오디오 프로세서, 시스템, 방법 및 컴퓨터 프로그램
CN116723438A (zh) * 2023-05-26 2023-09-08 三星电子(中国)研发中心 修正参数生成方法和装置

Similar Documents

Publication Publication Date Title
WO2020184842A1 (fr) Dispositif électronique et son procédé de commande
WO2019107868A1 (fr) Appareil et procédé de sortie de signal audio, et appareil d'affichage l'utilisant
WO2010087630A2 (fr) Procédé et appareil pour décoder un signal audio
WO2018056624A1 (fr) Dispositif électronique et procédé de commande associé
WO2018088742A1 (fr) Appareil d'affichage et son procédé de commande
WO2019078617A1 (fr) Appareil électronique et procédé de reconnaissance vocale
WO2014088328A1 (fr) Appareil de fourniture audio et procédé de fourniture audio
WO2018182274A1 (fr) Procédé et dispositif de traitement de signal audio
WO2015156654A1 (fr) Procédé et appareil permettant de représenter un signal sonore, et support d'enregistrement lisible par ordinateur
WO2018147701A1 (fr) Procédé et appareil conçus pour le traitement d'un signal audio
WO2021096233A1 (fr) Appareil électronique et son procédé de commande
WO2020050609A1 (fr) Appareil d'affichage et procédé de commande associé
WO2010008229A1 (fr) Appareil de codage et de décodage audio multi-objet prenant en charge un signal post-sous-mixage
WO2020050508A1 (fr) Appareil d'affichage d'image et son procédé de fonctionnement
WO2021010549A1 (fr) Appareil d'affichage et son procédé de commande
WO2015147619A1 (fr) Procédé et appareil pour restituer un signal acoustique, et support lisible par ordinateur
WO2010087631A2 (fr) Procédé et appareil pour décoder un signal audio
WO2016114432A1 (fr) Procédé de traitement de sons sur la base d'informations d'image, et dispositif correspondant
WO2022059869A1 (fr) Dispositif et procédé pour améliorer la qualité sonore d'une vidéo
WO2020226390A1 (fr) Dispositif électronique et son procédé de commande
WO2022039310A1 (fr) Terminal et procédé de délivrance de données audio multicanal au moyen d'une pluralité de dispositifs audio
EP3797346A1 (fr) Appareil d'affichage et procédé de commande associé
WO2019199040A1 (fr) Procédé et dispositif de traitement d'un signal audio, utilisant des métadonnées
EP4007949A1 (fr) Appareil d'affichage et son procédé de commande
WO2025234858A1 (fr) Dispositif électronique et procédé de traitement de signal audio

Legal Events

Date Code Title Description
121 Ep: the epo has been informed by wipo that ep was designated in this application

Ref document number: 25810293

Country of ref document: EP

Kind code of ref document: A1