Prosecution Insights
Last updated: August 15, 2026
Application No. 18/844,368

METHODS, APPARATUS, AND SYSTEMS FOR PROCESSING AUDIO SCENES FOR AUDIO RENDERING

Non-Final OA §102§103
Filed
Sep 05, 2024
Priority
Mar 09, 2022 — provisional 63/318,080 +2 more
Examiner
SAUNDERS JR, JOSEPH
Art Unit
2692
Tech Center
2600 — Communications
Assignee
Dolby International AB
OA Round
1 (Non-Final)
73%
Grant Probability
Favorable
1-2
OA Rounds
10m
Est. Remaining
94%
With Interview

Examiner Intelligence

Grants 73% — above average
73%
Career Allowance Rate
555 granted / 759 resolved
+11.1% vs TC avg
Strong +20% interview lift
Without
With
+20.5%
Interview Lift
resolved cases with interview
Typical timeline
2y 10m
Avg Prosecution
28 currently pending
Career history
782
Total Applications
across all art units

Statute-Specific Performance

§101
5.9%
-34.1% vs TC avg
§103
42.5%
+2.5% vs TC avg
§102
26.9%
-13.1% vs TC avg
§112
14.7%
-25.3% vs TC avg
Black line = Tech Center average estimate • Based on career data from 759 resolved cases

Office Action

§102 §103
DETAILED ACTION Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . This Office action is based on the communications filed September 5, 2024. Claims 32 – 54 are currently pending and considered below. Information Disclosure Statement The information disclosure statement (IDS) submitted on September 5, 2024, IDS submitted on December 10, 2024, and IDS submitted on August 18, 2025 is in compliance with the provisions of 37 CFR 1.97. Accordingly, the information disclosure statement is being considered by the examiner. Claim Rejections - 35 USC § 102 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action: A person shall be entitled to a patent unless – (a)(1) the claimed invention was patented, described in a printed publication, or in public use, on sale, or otherwise available to the public before the effective filing date of the claimed invention. (a)(2) the claimed invention was described in a patent issued under section 151, or in an application for patent published or deemed published under section 122(b), in which the patent or application, as the case may be, names another inventor and was effectively filed before the effective filing date of the claimed invention. Claim(s) 32 – 37 and 39 – 54 is/are rejected under 35 U.S.C. 102(a)(1) and 35 U.S.C. 102(a)(2) as being anticipated by Hammershoi et al. (US 2016/0109284 A1), hereinafter Hammershoi. Claim 32: Hammershoi discloses a method of processing audio scene information for audio rendering, the method comprising: receiving an audio scene description, the audio scene description comprising a representation of a three-dimensional audio scene (see at least, "FIG. 1 illustrates an embodiment where a depth camera, e.g. Kinect™device, provides data that allows a 3D cloud point data representation of the room. The whole procedure from a room scanning to the sound transmission module of a room acoustics modelling software comprises two steps. First, the room interior is scanned with the depth camera (Kinect™) using 3D scanning software. Then, the 3D point cloud of the room is processed using the Point Cloud Library (PCL) [www.pointcloud.org], in order to make an optimal room geometrical model for the sound transmission calculation,” Hammershoi [0053], “The proposed method is independent of the input scanning device ("depth camera"). In general, all scanning devices that provide a boundary description of a scanned room interior can be used to obtain a 3D point cloud model,” Hammershoi [0054]) and information on a source location of a sound source within the audio scene (see at least, “The raw scanning data SCD from the camera 3 _ CM is applied to a processor P which is prorammed to perform a suitable processing, such as explained in connection with the method embodiments, i.e. to generate a numerical representation of the room in response to the geometrical data, to apply an acoustical property to elements of the numerical representation of the room to obtain an acoustical room model, and to calculate an output indicative of the acoustical sound transmission in the room in response to the acoustical room model,” Hammershoi [0056], “In the embodiment of FIG. 2, the processor P is programmed to calculate a room impulse response RIR as output in the form of data representing a list of acoustic waves incident from a source position at a receiver position, including information regarding incidence angle and arrival time,” Hammershoi [0057]), wherein the representation of the three-dimensional audio scene is a voxel-based representation and comprises one or more indications of cuboid volumes in a voxel grid (see at least, “FIG. 5 illustrates a granular geometry representation obtained by the second approach, thus providing an input for e.g. a discrete ray tracing simulation algorithm. A voxel grid of a room is created and filled with the obtained point cloud from the depth-camera. Each voxel is defined as an occupied (contains point cloud data), or empty (does not contain point cloud data). Depending on the voxel type, different properties are assigned in order to define the behaviour of a sound ray when it reaches the voxel, e.g. reflection, diffraction, scattering, transmission etc. The voxel grid resolution defines the wanted level of details. A grid can be defined as a uniform grid when all space units are cubes of the same size but also as a hierarchical structure (k-d tree or octree ). The latter one improves efficiency of the sound transmission module by speeding up the ray traversal algorithm,” Hammershoi [0088]), and wherein each such indication comprises information on a pair of extreme-corner voxels of the cuboid volume (see at least, “1) Two points are detected from the point cloud model: min and max. Min is defined as minx, min y and min z coordinates from the whole point cloud. Max is defined as max x, max y and max z from the whole point cloud,” Hammershoi [0095], “2) A boundary box is created using min as a bottom left comer and max as a top right corner. It is divided into voxels giving the uniform voxel grid. Dimension of the voxel represents a resolution of the grid. FIG. 1 shows a boundary box filled with the point cloud model,” Hammershoi [0096]) and information on a common voxel property of the voxels in the cuboid volume (see at least, “3) Each voxel is indexed and it's content is been examined: if the voxel contains points of the point cloud model, it is labelled as "occupied" if it doesn't contain any point it's labelled as "empty",” Hammershoi [0097], “FIG. 7 shows steps of a method embodiment. The first step is to receive geometrical data from a 3D scanning device RDC in the form of raw measured data from the scanning device, or in a processed format, after having scanned a room. Next, a voxel representation is generated GVR in response to the received input data. Then, an acoustical property AAP is applied to the voxel elements of the voxel representation. In a simple version, a voxel is applied with the value "1 ", if it represents an acoustically reflecting voxel, and "0" if it represent air, thus defining a simple voxel-based acoustical room model,” Hammershoi [0102]); receiving an indication of a listener location of a listener within the audio scene (see at least, “FIG. 2 illustrates a block diagram of a system embodiment, where the described method is used to provide a binaural output signal L, R thus allowing a user to have the acoustic input allowing the user to listen to artificially generated sound creating the feeling of being present in a room which has been scanned with a 3D camera 3_CM, e.g. the Kinect™ device,” [0056], “This data RIR is then applied to a binaural processor BP which receives the room impulse response data RIR, as well as a sound signal input S_I, and a stream of data indicating an orientation O_D of the user, e.g. also the user's 3D position within the room. Thus, such system allows tracking of the user moving around in the room, and thereby capable of generating binaural signals L, R to the user corresponding to the left and right ear signals, if the user was present in a location within the room, and with sound S_I from a receiver position within the room,” Hammershoi [0057]); obtaining diffraction information relating to an acoustic diffraction path within the audio scene between the source location and the listener location (see at least, “FIG. 5 illustrates a granular geometry representation obtained by the second approach, thus providing an input for e.g. a discrete ray tracing simulation algorithm. A voxel grid of a room is created and filled with the obtained point cloud from the depth-camera. Each voxel is defined as an occupied (contains point cloud data), or empty (does not contain point cloud data). Depending on the voxel type, different properties are assigned in order to define the behaviour of a sound ray when it reaches the voxel, e.g. reflection, diffraction, scattering, transmission etc. The voxel grid resolution defines the wanted level of details. A grid can be defined as a uniform grid when all space units are cubes of the same size but also as a hierarchical structure (k-d tree or octree ). The latter one improves efficiency of the sound transmission module by speeding up the ray traversal algorithm,” Hammershoi [0088]); performing audio rendering for the sound source based on the diffraction information (see at least, “The binaural processor BP applies Head-Related Transfer Functions (HRTF s) to the sound input S_I in accordance with the RIR data, i.e. left and right outputs are generated in response to each incoming sound wave by applying appropriate HRTFs corresponding to the angle of incidence of the sound wave, taking into account the orientation O D of the user relative to the room. With e.g. a discrete ray tracing method implemented on the processor P, such system shown in FIG. 2 can be used for real-time rendering applications, e.g. virtual reality, tele presence etc. The processor P may be implemented as a Personal Computer, or a dedicated device, or a combination thereof,” Hammershoi [0057]); and outputting a representation of the diffraction information (see at least, “To sum up, the invention provides a method for generating an output indicative of acoustical sound transmission in a room. By using e.g. a point cloud representation of an acoustic environment, it is possible to calculate its acoustics from the interior information obtained from the depth camera. This approach is suitable e.g. for run-time applications since it is not based on an audible excitation that can disturb running audio. Also, the point-cloud model can be updated in real time according to the scene changes detected by depth-camera. This allows efficient acoustical simulation of dynamic, interactive environments. Although only geometrical information of a room is provided, high amount of surface details leaves possibility for implementation of material recognition algorithms that involve semantic mapping. This can provide information of reflective properties of surfaces or objects at a point level. Also, a high amount of details allows a good approximation of complex geometries, e.g. porous materials, and rough surfaces, thus a more natural simulation of wave phenomena like diffraction and scattering is possible,” Hammershoi [0103]). Claim 33: Hammershoi discloses the method according to claim 32, wherein outputting the representation of the diffraction information comprises outputting a data element comprising the diffraction information and information on a scene state, the scene state comprising the audio scene description and the listener location (see at least, “Thus, in one scenario the geometrical scanning is performed once, but in other scenario the scanning device provides a stream of geometrical data for dynamic processing, thus allowing detection of moving objects or other changes influencing sound transmission properties of the room, e.g. opening of a door or window in the room,” Hammershoi [0011]], This data RIR is then applied to a binaural processor BP which receives the room impulse response data RIR, as well as a sound signal input S_I, and a stream of data indicating an orientation O_D of the user, e.g. also the user's 3D position within the room. Thus, such system allows tracking of the user moving around in the room, and thereby capable of generating binaural signals L, R to the user corresponding to the left and right ear signals, if the user was present in a location within the room, and with sound S_I from a receiver position within the room,” Hammershoi [0057], “FIG. 5 illustrates a granular geometry representation obtained by the second approach, thus providing an input for e.g. a discrete ray tracing simulation algorithm. A voxel grid of a room is created and filled with the obtained point cloud from the depth-camera. Each voxel is defined as an occupied (contains point cloud data), or empty (does not contain point cloud data). Depending on the voxel type, different properties are assigned in order to define the behaviour of a sound ray when it reaches the voxel, e.g. reflection, diffraction, scattering, transmission etc. The voxel grid resolution defines the wanted level of details. A grid can be defined as a uniform grid when all space units are cubes of the same size but also as a hierarchical structure (k-d tree or octree ). The latter one improves efficiency of the sound transmission module by speeding up the ray traversal algorithm,” Hammershoi [0088]) Claim 34: Hammershoi discloses the method according to claim 32, wherein the representation of the diffraction information is output to a bitstream (see at least, “Thus, in one scenario the geometrical scanning is performed once, but in other scenario the scanning device provides a stream of geometrical data for dynamic processing, thus allowing detection of moving objects or other changes influencing sound transmission properties of the room, e.g. opening of a door or window in the room,” Hammershoi [0011], “This data RIR is then applied to a binaural processor BP which receives the room impulse response data RIR, as well as a sound signal input S_I, and a stream of data indicating an orientation O_D of the user, e.g. also the user's 3D position within the room. Thus, such system allows tracking of the user moving around in the room, and thereby capable of generating binaural signals L, R to the user corresponding to the left and right ear signals, if the user was present in a location within the room, and with sound S_I from a receiver position within the room,” Hammershoi [0057], “6) According to the voxel labels, a .txt file is created that can be directly used as an input for the sound transmission module,” Hammershoi [0100]). Claim 35: Hammershoi discloses the method according to claim 32, wherein the diffraction information is output for later re-use for audio rendering by the same rendering instance or for later re-use by another rendering instance (see at least, “Also, an advantage of a point-based geometry representation is an ability to store information about acous-tical properties of each point directly in the 3D point-cloud model just as another property of the PLY header. Eventually, this can facilitate calculation of the sound transmission within the room,” Hammershoi [0085]). Claim 36: Hammershoi discloses the method according to claim 32, wherein the information on the pair of extreme-corner voxels of the cuboid volume comprises indications of respective voxel indices assigned to the extreme-corner voxels, the voxels of the voxel-based audio scene representation having uniquely assigned consecutive voxel indices (see at least, “1) Two points are detected from the point cloud model: min and max. Min is defined as minx, min y and min z coordinates from the whole point cloud. Max is defined as max x, max y and max z from the whole point cloud,” Hammershoi [0095], “2) A boundary box is created using min as a bottom left comer and max as a top right corner. It is divided into voxels giving the uniform voxel grid. Dimension of the voxel represents a resolution of the grid. FIG. 1 shows a boundary box filled with the point cloud model,” Hammershoi [0096]). Claim 37: Hammershoi discloses the method according to claim 32, wherein the representation of the three- dimensional audio scene is a voxel-based representation; wherein the diffraction information comprises an indication of a location of a voxel that is located on or on the proximity of the diffraction path and an indication of a length of the diffraction path; and wherein the indication of the location of the voxel located on or on the proximity of the diffraction path is an indication of a voxel index assigned to said voxel, the voxels of the voxel-based audio scene representation having uniquely assigned consecutive voxel indices (see at least, “5) Each occupied voxel is examined by their neighbour voxels in order to determine the reflection type after a ray reaches it. If the voxel belongs to the group of voxels that represents a surface parallel with the x-axis it is labelled with X. The same approach is for Y and Z voxels. If the voxel has two or three labels at the same time (XY, YZ, XZ or XYZ) it is labelled as a corner-C. These labels allow to define specific interaction between ray and the voxel-the way how the ray reflects,” Hammershoi [0099]). Claim 39: Hammershoi discloses the method according to claim 32, further comprising: if it is determined that the current scene state does not correspond to a known scene state, determining the diffraction information using a pathfinding algorithm, based on the source location, the listener location, and the representation of the three-dimensional audio scene (see at least, “Finally, a ray tracing algorithm is applied for calculating acoustical sound transmission in the room, e.g. form a source position to a receiver position, so as to generate a ray tracing output GRT_O based on the voxelbased acoustical room model,” Hammershoi [0102]). Claim 40: Hammershoi discloses the method according to claim 32, wherein the representation of the three- dimensional audio scene is a voxel-based representation including information on locations and material properties of a plurality of occluder voxels (see at least, “By "photo" is understood, e.g. in connection with the Kinect™, RGB information of each point which provides a texture of a surface that can be used for pattern recognition in order to predict a scanned material and thus its acoustical properties. E.g. colors and patterns in the photo may be used to predict acoustical properties of a given surface or object, such as acoustical absorption, thus allowing automatically assigning acoustical properties to elements of the numerical representation of the room accordingly. Especially, acoustical absorption data may be assigned to an object in accordance with a database of acoustical absorption data. In a simple embodiment, the geometrical data and/or a photo is processed with respect to identifying surfaces or objects with similar material, determining which type of material it is by selecting between a number of predetermined definitions materials, and assigning prestored acoustical properties for that type of material to geometrical elements representing the object in the model,” Hammershoi [0018], “4) Empty voxels represent air-ray can go through them with no interaction while occupied voxels represent surfaces that ray reflects from,” Hammershoi [0098]). Claim 41: Hammershoi discloses a method of compressing an audio scene for three-dimensional audio rendering, the method comprising: obtaining a voxelized representation of the audio scene, the voxelized representation comprising a plurality of voxels arranged in a voxel grid, each voxel having an associated voxel property (see at least, “FIG. 5 illustrates a granular geometry representation obtained by the second approach, thus providing an input for e.g. a discrete ray tracing simulation algorithm. A voxel grid of a room is created and filled with the obtained point cloud from the depth-camera. Each voxel is defined as an occupied (contains point cloud data), or empty (does not contain point cloud data). Depending on the voxel type, different properties are assigned in order to define the behaviour of a sound ray when it reaches the voxel, e.g. reflection, diffraction, scattering, transmission etc. The voxel grid resolution defines the wanted level of details. A grid can be defined as a uniform grid when all space units are cubes of the same size but also as a hierarchical structure (k-d tree or octree ). The latter one improves efficiency of the sound transmission module by speeding up the ray traversal algorithm,” Hammershoi [0088]); determining, among the voxels of the voxelized representation, a set of voxels that forms a connected geometric region on the voxel grid (see at least, “By maintaining a list of points from which the mesh can be grown and extending it until all possible points are connected [3D scanning software, http://skanect.manctl.com/], a concave hull can be created that represents room boundaries and the inner surfaces,” Hammershoi [0086], “FIG. 5 illustrates a granular geometry representation obtained by the second approach, thus providing an input for e.g. a discrete ray tracing simulation algorithm. A voxel grid of a room is created and filled with the obtained point cloud from the depth-camera. Each voxel is defined as an occupied (contains point cloud data), or empty (does not contain point cloud data). Depending on the voxel type, different properties are assigned in order to define the behaviour of a sound ray when it reaches the voxel, e.g. reflection, diffraction, scattering, transmission etc. The voxel grid resolution defines the wanted level of details. A grid can be defined as a uniform grid when all space units are cubes of the same size but also as a hierarchical structure (k-d tree or octree ). The latter one improves efficiency of the sound transmission module by speeding up the ray traversal algorithm,” Hammershoi [0088], “5) Each occupied voxel is examined by their neighbour voxels in order to determine the reflection type after a ray reaches it. If the voxel belongs to the group of voxels that represents a surface parallel with the x-axis it is labelled with X. The same approach is for Y and Z voxels. If the voxel has two or three labels at the same time (XY, YZ, XZ or XYZ) it is labelled as a corner-C. These labels allow to define specific interaction between ray and the voxel-the way how the ray reflects,” Hammershoi [0099]), wherein the geometric region has a cuboid shape (see at least, “1) Two points are detected from the point cloud model: min and max. Min is defined as minx, min y and min z coordinates from the whole point cloud. Max is defined as max x, max y and max z from the whole point cloud,” Hammershoi [0095], “2) A boundary box is created using min as a bottom left comer and max as a top right corner. It is divided into voxels giving the uniform voxel grid. Dimension of the voxel represents a resolution of the grid. FIG. 1 shows a boundary box filled with the point cloud model,” Hammershoi [0096]) and the voxels in the geometric region share a common voxel property (see at least, “3) Each voxel is indexed and it's content is been examined: if the voxel contains points of the point cloud model, it is labelled as "occupied" if it doesn't contain any point it's labelled as "empty",” Hammershoi [0097], “FIG. 7 shows steps of a method embodiment. The first step is to receive geometrical data from a 3D scanning device RDC in the form of raw measured data from the scanning device, or in a processed format, after having scanned a room. Next, a voxel representation is generated GVR in response to the received input data. Then, an acoustical property AAP is applied to the voxel elements of the voxel representation. In a simple version, a voxel is applied with the value "1 ", if it represents an acoustically reflecting voxel, and "0" if it represent air, thus defining a simple voxel-based acoustical room model,” Hammershoi [0102]); determining, from the plurality of voxels of the voxelized representation, at least a first boundary voxel and a second boundary voxel for the set of voxels, the first boundary voxel and the second boundary voxel defining the cuboid shape of the geometric region (see at least, “1) Two points are detected from the point cloud model: min and max. Min is defined as minx, min y and min z coordinates from the whole point cloud. Max is defined as max x, max y and max z from the whole point cloud,” Hammershoi [0095], “2) A boundary box is created using min as a bottom left comer and max as a top right corner. It is divided into voxels giving the uniform voxel grid. Dimension of the voxel represents a resolution of the grid. FIG. 1 shows a boundary box filled with the point cloud model,” Hammershoi [0096]); and generating a representation of the audio scene based on the determined set of voxels (see at least, "FIG. 1 illustrates an embodiment where a depth camera, e.g. Kinect™device, provides data that allows a 3D cloud point data representation of the room. The whole procedure from a room scanning to the sound transmission module of a room acoustics modelling software comprises two steps. First, the room interior is scanned with the depth camera (Kinect™) using 3D scanning software. Then, the 3D point cloud of the room is processed using the Point Cloud Library (PCL) [www.pointcloud.org], in order to make an optimal room geometrical model for the sound transmission calculation,” Hammershoi [0053], “The proposed method is independent of the input scanning device ("depth camera"). In general, all scanning devices that provide a boundary description of a scanned room interior can be used to obtain a 3D point cloud model,” Hammershoi [0054]) and information on a source location of a sound source within the audio scene (see at least, “The raw scanning data SCD from the camera 3 _ CM is applied to a processor P which is prorammed to perform a suitable processing, such as explained in connection with the method embodiments, i.e. to generate a numerical representation of the room in response to the geometrical data, to apply an acoustical property to elements of the numerical representation of the room to obtain an acoustical room model, and to calculate an output indicative of the acoustical sound transmission in the room in response to the acoustical room model,” Hammershoi [0056], “In the embodiment of FIG. 2, the processor P is programmed to calculate a room impulse response RIR as output in the form of data representing a list of acoustic waves incident from a source position at a receiver position, including information regarding incidence angle and arrival time,” Hammershoi [0057], “To sum up, the invention provides a method for generating an output indicative of acoustical sound transmission in a room. By using e.g. a point cloud representation of an acoustic environment, it is possible to calculate its acoustics from the interior information obtained from the depth camera. This approach is suitable e.g. for run-time applications since it is not based on an audible excitation that can disturb running audio. Also, the point-cloud model can be updated in real time according to the scene changes detected by depth-camera. This allows efficient acoustical simulation of dynamic, interactive environments. Although only geometrical information of a room is provided, high amount of surface details leaves possibility for implementation of material recognition algorithms that involve semantic mapping. This can provide information of reflective properties of surfaces or objects at a point level. Also, a high amount of details allows a good approximation of complex geometries, e.g. porous materials, and rough surfaces, thus a more natural simulation of wave phenomena like diffraction and scattering is possible,” Hammershoi [0103]). Claim 42: Hammershoi discloses the method of claim 41, further comprising determining, for the geometric region, at least one scene element parameter comprising one or more of: a scene element identifier, an acoustic property identifier and/or audio rendering instruction set identifier, and indices of the corresponding first and second boundary voxels defining the geometric region (see at least, “2) A boundary box is created using min as a bottom left comer and max as a top right corner. It is divided into voxels giving the uniform voxel grid. Dimension of the voxel represents a resolution of the grid. FIG. 1 shows a boundary box filled with the point cloud model,” Hammershoi [0096], “3) Each voxel is indexed and it's content is been examined: if the voxel contains points of the point cloud model, it is labelled as "occupied" if it doesn't contain any point it's labelled as "empty",” Hammershoi [0097]). Claim 43: Hammershoi discloses the method of claim 42, further comprising applying entropy coding to the at least one scene element parameter for the geometric region (see at least, “FIG. 5 illustrates a granular geometry representation obtained by the second approach, thus providing an input for e.g. a discrete ray tracing simulation algorithm. A voxel grid of a room is created and filled with the obtained point cloud from the depth-camera. Each voxel is defined as an occupied (contains point cloud data), or empty (does not contain point cloud data). Depending on the voxel type, different properties are assigned in order to define the behaviour of a sound ray when it reaches the voxel, e.g. reflection, diffraction, scattering, transmission etc. The voxel grid resolution defines the wanted level of details. A grid can be defined as a uniform grid when all space units are cubes of the same size but also as a hierarchical structure (k-d tree or octree ). The latter one improves efficiency of the sound transmission module by speeding up the ray traversal algorithm,” Hammershoi [0088], “The result of Step 1 is a point cloud model of the room represented by .ply file format. Density of the points is very high (30202895 points in total) and the file itself is very heavy (1.3 GB of row text file),” Hammershoi [0092], “In Step 2, the obtained .ply file is processed in order to reduce the amount of data. Using PCL and its voxel_grid library, the number of points is reduced to 17896. It has been done by octree structure with the leaf size of 0.1 m,” Hammershoi [0093]). Claim 44: Hammershoi discloses the method claim 42, further comprising outputting a bitstream including the at least one scene element parameter for determining the set of voxels associated with the geometric region for a compressed representation of the audio scene based on the determined set of voxels (see at least, “Thus, in one scenario the geometrical scanning is performed once, but in other scenario the scanning device provides a stream of geometrical data for dynamic processing, thus allowing detection of moving objects or other changes influencing sound transmission properties of the room, e.g. opening of a door or window in the room,” Hammershoi [0011], “This data RIR is then applied to a binaural processor BP which receives the room impulse response data RIR, as well as a sound signal input S_I, and a stream of data indicating an orientation O_D of the user, e.g. also the user's 3D position within the room. Thus, such system allows tracking of the user moving around in the room, and thereby capable of generating binaural signals L, R to the user corresponding to the left and right ear signals, if the user was present in a location within the room, and with sound S_I from a receiver position within the room,” Hammershoi [0057], “6) According to the voxel labels, a .txt file is created that can be directly used as an input for the sound transmission module,” Hammershoi [0100]). Claim 45: Hammershoi discloses the method according to claim 41, wherein the geometric region is related to a scene element within the audio scene (see at least, “2) A boundary box is created using min as a bottom left comer and max as a top right corner. It is divided into voxels giving the uniform voxel grid. Dimension of the voxel represents a resolution of the grid. FIG. 1 shows a boundary box filled with the point cloud model,” Hammershoi [0096], “3) Each voxel is indexed and it's content is been examined: if the voxel contains points of the point cloud model, it is labelled as "occupied" if it doesn't contain any point it's labelled as "empty",” Hammershoi [0097]). Claim 46: Hammershoi discloses the method according to claim 41, wherein the audio scene comprises a large scene represented by the determined set of voxels, the large scene including a set of sub-scenes, wherein each of the sub-scenes corresponds to a subset of the determined set of voxels, the method further comprising determining, among the determined set of voxels, the subsets of voxels for the corresponding sub-scenes (see at least, “The Kinect™ was placed in the middle of the room and aligned with the room edges. The coordinates of all obtained points are relative to the device's starting position and orientation and the alignment has to be done in order to relate point cloud with the room coordinate system (length of the room-x axis, width-y axis and hight-z axis). Another option is to know in advance the position of the scanning device and orientation at the starting point, and than to correct the whole set of obtained data with this offset. When the first frame is recorded, the device is rotated by z and y axes in order to cover room interior as much as possible. While rotating the Kinect, the software maps new acquired scene to the previous one detecting "key points" in both scenes and aligning them at the same 3D model. The quality of the resulting model depends very much on this step-if the scenes are not matched in a good way, a flat surface can be represented in a "broken" or multiple planes which introduces errors for later processing and calculations,” Hammershoi [0092]). Claim 47: Hammershoi discloses the method according to claim 41, further comprising applying interpolation of audio voxels in time and/or space (see at least, “The Kinect output is post-processed in order to obtain a 3D point-cloud model of the room. Different post processing is applied for the optimal 3D point-cloud models to be used as an input for a sound transmission module of a different physically based room acoustics modelling methods. In general, granular representation of a room's inner surfaces makes it possible to have a high level of details when creating a geometry model. A high density of points makes it possible to recognize fine details and sharp edges of the surfaces as well as transitions between adjacent areas. Using a different resolution for a room model, it is possible to take into account any frequency-dependent geometry. Thus, on the basis of the original model, a downsampled point-cloud models can be created with less details, e.g. for low frequency simulations,” Hammershoi [0085], “The described method is capable of providing fast acquisition of the arbitrary room interior, not just the boundaries but also the inner surfaces. It is scalable and by changing the resolution of the uniform voxel grid high level of details can be provided, thus allowing application of advanced acoustic properties to the surface, thereby providing a realistic acoustic room model. The process is autonomous and the input to the sound transmission module can be generated directly from the room scan,” Hammershoi [0101]). Claim 48: Hammershoi discloses the method according to claim 41, further comprising redefining voxel properties for a subset of the set of voxels associated with a scene sub-element in the geometric region for overwriting the subset with the redefined voxel properties (see at least, “One embodiment comprises performing a discrete ray tracing algorithm in response to the acoustical room model. As mentioned above, especially in combination with a voxel-based geometrical representation, discrete ray tracing is suitable for dynamically updating the acoustical sound transmission properties in the model of the room. Thus, especially the method may comprise receiving a stream of geometrical data representing respective scannings of the room, and dynamically updating a discrete ray tracing representation accordingly,” Hammershoi [0024]). Claim 49: Hammershoi discloses the method according to claim 41, further comprising determining a superset of voxels including the determined set of voxels, the determined set of voxels associated with a scene sub-element within the geometric region, the method further comprising assigning a new voxel property to the determined set of voxels and overwriting the voxel property of the determined set of voxels with the new voxel property (see at least, “The Kinect™ was placed in the middle of the room and aligned with the room edges. The coordinates of all obtained points are relative to the device's starting position and orientation and the alignment has to be done in order to relate point cloud with the room coordinate system (length of the room-x axis, width-y axis and hight-z axis). Another option is to know in advance the position of the scanning device and orientation at the starting point, and than to correct the whole set of obtained data with this offset. When the first frame is recorded, the device is rotated by z and y axes in order to cover room interior as much as possible. While rotating the Kinect, the software maps new acquired scene to the previous one detecting "key points" in both scenes and aligning them at the same 3D model. The quality of the resulting model depends very much on this step-if the scenes are not matched in a good way, a flat surface can be represented in a "broken" or multiple planes which introduces errors for later processing and calculations,” [0092], “One embodiment comprises performing a discrete ray tracing algorithm in response to the acoustical room model. As mentioned above, especially in combination with a voxel-based geometrical representation, discrete ray tracing is suitable for dynamically updating the acoustical sound transmission properties in the model of the room. Thus, especially the method may comprise receiving a stream of geometrical data representing respective scannings of the room, and dynamically updating a discrete ray tracing representation accordingly,” Hammershoi [0024]). Claim 50: Hammershoi discloses the method according to claim 41, further comprising determining a voxel size for representing the geometric region, wherein the voxel size is based on a number of voxels along a scene dimension of the geometric region (see at least, “FIG. 5 illustrates a granular geometry representation obtained by the second approach, thus providing an input for e.g. a discrete ray tracing simulation algorithm. A voxel grid of a room is created and filled with the obtained point cloud from the depth-camera. Each voxel is defined as an occupied (contains point cloud data), or empty (does not contain point cloud data). Depending on the voxel type, different properties are assigned in order to define the behaviour of a sound ray when it reaches the voxel, e.g. reflection, diffraction, scattering, transmission etc. The voxel grid resolution defines the wanted level of details. A grid can be defined as a uniform grid when all space units are cubes of the same size but also as a hierarchical structure (k-d tree or octree ). The latter one improves efficiency of the sound transmission module by speeding up the ray traversal algorithm,” Hammershoi [0088]). Claim 51: Hammershoi discloses an apparatus comprising a processor and a memory coupled to the processor, and storing instructions for the processor, wherein the processor is adapted to carry out the method according to claim 32 (see at least, “FIG. 2 illustrates a block diagram of a system embodiment, where the described method is used to provide a binaural output signal L, R thus allowing a user to have the acoustic input allowing the user to listen to artificially generated sound creating the feeling of being present in a room which has been scanned with a 3D camera 3_CM, e.g. the Kinect™ device. The raw scanning data SCD from the camera 3 _ CM is applied to a processor P which is prorammed to perform a suitable processing, such as explained in connection with the method embodiments, i.e. to generate a numerical representation of the room in response to the geometrical data, to apply an acoustical property to elements of the numerical representation of the room to obtain an acoustical room model, and to calculate an output indicative of the acoustical sound transmission in the room in response to the acoustical room model,” Hammershoi [0056], “In a fourth aspect, the invention provides a non-transitory, computer readable storage medium with a computer executable program code adapted to perform the method according to the first aspect,” Hammershoi [0037]). Claim 52: Hammershoi discloses an apparatus , comprising a processor and a memory coupled to the processor, and storing instructions for the processor, wherein the processor is adapted to carry out the method according to claim 41 (see at least, “FIG. 2 illustrates a block diagram of a system embodiment, where the described method is used to provide a binaural output signal L, R thus allowing a user to have the acoustic input allowing the user to listen to artificially generated sound creating the feeling of being present in a room which has been scanned with a 3D camera 3_CM, e.g. the Kinect™ device. The raw scanning data SCD from the camera 3 _ CM is applied to a processor P which is prorammed to perform a suitable processing, such as explained in connection with the method embodiments, i.e. to generate a numerical representation of the room in response to the geometrical data, to apply an acoustical property to elements of the numerical representation of the room to obtain an acoustical room model, and to calculate an output indicative of the acoustical sound transmission in the room in response to the acoustical room model,” Hammershoi [0056], “In a fourth aspect, the invention provides a non-transitory, computer readable storage medium with a computer executable program code adapted to perform the method according to the first aspect,” Hammershoi [0037]). Claim 53: Hammershoi discloses a non-transitory computer program comprising instructions that, when executed by a processor, cause the processor to carry out the method according to claim 32 (see at least, “In a fourth aspect, the invention provides a non-transitory, computer readable storage medium with a computer executable program code adapted to perform the method according to the first aspect,” Hammershoi [0037]). Claim 54: Hammershoi discloses a non-transitory computer program comprising instructions that, when executed by a processor, cause the processor to carry out the method according to claim 41 (see at least, “In a fourth aspect, the invention provides a non-transitory, computer readable storage medium with a computer executable program code adapted to perform the method according to the first aspect,” Hammershoi [0037]). Claim Rejections - 35 USC § 103 The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claim(s) 38 is/are rejected under 35 U.S.C. 103 as being unpatentable over Hammershoi in view of Schissler et al. (US 2015/0378019 A1), hereinafter Schissler. Claim 38: Hammershoi discloses the method according to claim 32, but does not disclose wherein determining whether the current scene state corresponds to a known scene state comprises determining a hash value based on the current scene state. However, Schissler discloses a similar invention directed to interactive diffuse reflections and higher-order diffraction in virtual environment scenes. Schissler further discloses wherein determining whether the current scene state corresponds to a known scene state comprises determining a hash value based on the current scene state (see at least, “In some embodiments, SPT module 108 may be configured to preserve the phase information of the sound rays in each of the aforementioned path tracing groups. For example, SPT module 108 may determine, for each of the path tracing groups, a sound delay (e.g., an total sound delay) by combining a sound delay computed for the current time with one or more previously computed reflected sound delays respectively associated with the previously elapsed times. In some embodiments, the one or more previously computed reflected sound intensities and the one or more previously computed reflected sound delays may each comprise a moving average. In some embodiments, SPT module 108 may store the determined sound intensity for each of the path tracing groups within an entry of a hash table cache. In one embodiment, the hash table cache may be stored in memory 104. Notably, each entry of the hash table cache may be repeatedly and/or periodically updated by SPT module 108, e.g., for each time frame segment of a time period associated with an acoustic simulation duration,” Schissler [0028]). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to incorporate the aforementioned features of Schissler in the invention of Hammershoi thereby allowing for the advantage of fast lookups of scene properties in the dynamically changing scene in invention of Hammershoi. Conclusion Any inquiry concerning this communication or earlier communications from the examiner should be directed to JOSEPH SAUNDERS whose telephone number is (571)270-1063. The examiner can normally be reached Monday-Thursday, 9:00 a.m. - 4 p.m., EST. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Carolyn R Edwards can be reached at (571)270-7136. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /JOSEPH SAUNDERS JR/Primary Examiner, Art Unit 2692
Read full office action

Prosecution Timeline

Sep 05, 2024
Application Filed
Jul 15, 2026
Non-Final Rejection mailed — §102, §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12705015
AUDIO PROCESSING SYSTEM AND METHOD
3y 2m to grant Granted Aug 11, 2026
Patent 12707186
WIRELESS HEADSET SYSTEM AND WIRELESS HEADSET
2y 9m to grant Granted Aug 11, 2026
Patent 12701379
Audio Scene Description and Control
3y 5m to grant Granted Aug 04, 2026
Patent 12699540
SYSTEMS AND METHODS FOR REDUCING AUDIO QUALITY BASED ON ACOUSTIC ENVIRONMENT
2y 10m to grant Granted Aug 04, 2026
Patent 12688003
SYSTEMS AND METHODS FOR SCALABLE MANAGEMENT OF AUDIO SYSTEM DEVICES
5y 0m to grant Granted Jul 21, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
73%
Grant Probability
94%
With Interview (+20.5%)
2y 10m (~10m remaining)
Median Time to Grant
Low
PTA Risk
Based on 759 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month