DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Response to Amendment
In response to the office action from 2/24/2026, the applicant has submitted an amendment, filed 5/22/2026, amending claims 1-2, 8, 10, 17, 19, adding claims cancelling claim 3, while arguing to traverse the prior art and 101 rejections. Applicant’s arguments have been fully considered and determined persuasive with respect to the 101 rejections based on the latest amendments, but are moot with respect to new grounds of rejections further in view of Herre et al. (WO 2019/012131) mandated by the latest amendments.
Response to Arguments
Pages 8 to the middle of page 9 provide a broad overview of the latest amendments and their support.
The examiner does acknowledge that the most recent amendments are supported.
Page 9 the paragraph before last discusses the previous claim objections.
Due to the latest amendments to claims 8 and 17, the said objections are withdrawn.
Page 9 the last paragraph discusses 101 rejection.
Due to the latest amendments, the said rejection is withdrawn.
The remaining part of the remarks is focused on the previous 102 and 103 rejections. It is concluded that no reference used in the first action teach the latest amendments.
The examiner agrees and requests visiting the new office action further in view of Herre et al.
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claim(s) 1-2, 4-7, 10-16, 19-21 is/are rejected under 35 U.S.C. 103 as being unpatentable over Bhatt et al. (US 2024/0185449), and further in view of Herre et al. (WO 2019/012131).
Regarding claim 1, Bhatt et al. do teach a device, the device comprising processing circuitry coupled to storage (Abstract S1: “A video conference call system” (a device) “is provided with a camera to generate an input frame image of a conference room”; ¶ 0081 lines 11+: “The information handling system may include random access memory (RAM), one or more processing resources such as a central processing unit (CPU) or hardware or software control logic”),
the processing circuitry configured to:
identify metadata comprising depth sensing information and camera information received from an in-room device located at a first location having a first camera (Abstract S1: “A video” (or “camera field of view” (metadata or camera information, ¶ 0035 2nd column line 9)) “conference call system is provided with a camera” (received in an in-room device) “to generate an input frame image of a conference room, where the video conference call system detects a human head”(detecting) “for each meeting participant” (one or more in-room users); ¶ 0002 S3: “detecting the location of meeting participant in a room location” “e.g., the distance” (including depth sensing information) “and direction between the camera and meeting participant” (associated with each in-room person); ¶ 0003 column 2 lines 7+: “identify” (identifying) “for each detected human head, a head bounding box with specified image plane coordinate and dimension information in a head box data structure” (a subset of the “video” (metadata)));
perform face recognition on one or more in-room users (Abstract lines 6-8: “generates a head bounding box which surround each detected human head” (for each face) “and identifies” (perform identification) “a corresponding meeting participant” (for each user of the one or more in-room users));
calculate a distance of a first in-room user based on the metadata and a first number of pixels across the face of the first in-room user (¶ 0048 S1: “determining the position and distance” (calculating a distance) “of a meeting participant” (of a first in-room user) “in a room based on the pixel count” (based on a first number of pixels) “for the height and width” (based on the “video” pictures (metadata)) “of each human head” (across his face));
calculate a distance between the first in-room user and a second in-room user based on the metadata and the first number of pixels across the face of the first in-room user and a second number of pixels across the face of the second in-room user (¶ 0078 first column last 3 lines+: “The disclosed methodology also applies the pixel width measure” (using the first number of pixels across each participant face) “and pixel height measure” “extracted from each head bounding box” (and metadata) “to one or more reverse lookup tables to extract meeting room coordinates” (determine distances of each in-room user from the in-room device because according to ¶ 0028 line 21 “coordinates [are] in relation to the camera location”) “for each meeting participant” (for e.g., a first in-room user as well as a second user, which enables calculating the distance between them, e.g., see Fig. 4 and ¶ 0031 lines 13+: “For example, a meeting participant 42” (e.g. a first in-room user) “located on the center line of sight of camera focal point” “at a distance of d=0.5 meters” “a meeting participant 43” (a second in-room user) “located on the center line of sight of camera” “at a distance of camera” “at a larger distance of d=1.0” (which readily gives the distance between “participant 42” (first in-room user) and “participant 43” (and second in-room user) as 0.5 meters)).
Bhatt et al. do not specifically disclose:
Generate a gain matrix for the one or more in-room users based, in part, on distances of the one or more in-room uses from the first camera and beamforming angles corresponding to the one or more in-room users; and
Apply gain values form the gain matrix to voice data corresponding to each of the one or more in-room users during post-processing on a downstream speaker playback path.
Herre et al. do teach:
Generate a gain matrix for the one or more in-room users based, in part, on distances of the one or more in-room uses from the first camera and beamforming angles corresponding to the one or more in-room users (page 27 lines 14+: “The distance-dependent gain” (gain depending on distance) “for each DoA is derived from the resulting length of the direction vector, dp(k,n)” (and direction vectors) and described as “Gi(rp(k,n))(||dp(k,n)||) -γ” (a gain matrix element dependent on “rp” (distance) and “dp” (direction), where “DoA” (“direction of arrival” (P.5 lines 3-4)) which “is derived” from “dp(k,n)” is according to page 9 lines 15-20 “calculate[ed]” “using” “meta data” which depends on “distance to a source” (distance from one or more users) “using two angles” (and beamforming angles) “with respect to two different reference locations and the distance positions”; page 4 lines 7-9: “arrays of cameras can be employed to generate light-field rendering” (e.g., using cameras to determine the distances and angles) “For audio, a similar set up employes distributed microphone arrays” (beamforming is used to help in audio “rendering”)) ; and
Apply gain values form the gain matrix to voice data corresponding to each of the one or more in-room users during post-processing on a downstream speaker playback path (page 3 lines 29-32: “The audio is often generated using object-based rendering” (during a post processing or playback of an “audio” (voice data)) “where each audio object is rendered with distance-dependent gain” (using the gain matrix gain values which depends on distance) “and relative direction” (and the beam forming angles) “from the user” (corresponding to one or more in room users) “based on the tracking data”).
It would have therefore been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to incorporate the “gain factor” incorporation into audio “rendering” (regeneration) of Herre et al. into the “sound processing” “to amplify the speech of [a] speaker who is located more than X meters form the camera” of Bhatt et al. (¶ 0027 page 3 lines 6-7) would enable the combined systems and their associated methods to perform in combination as they do separately and to further enable Bhatt et al. to achieve a “high-quality [audio] version” as disclosed in Herre et al. page 37 line 24 since Herre et al.’s “gain” depends on both “distance” and “direction”.
Regarding claim 2, Bhatt et al. do teach the device of claim 1, wherein the depth sensing information is associated with the one or more in-room users located in a field of view of the first camera and provides the distances of the one or more in-room users from the first camera (¶ 0002 S3: “detecting the location of meeting participant in a room location” “e.g., the distance” (the depth sensing information which provides distances from the “participants” (one or more in-room users) to the “camera” (the first camera)) “and direction between the camera and meeting participant” (associated with each in-room user being located in the field of view of the “camera” (the first camera))).
Regarding claim 4, Bhatt et al. do teach The device of claim 1, wherein the metadata information comprises reference frame information associated with each of the one or more in-room users captured at one or more time intervals (¶ 0078 first column last 3 lines+: “The disclosed methodology also applies the pixel width measure” “and pixel height measure” “extracted from each head bounding box” (the metadata) “to one or more reverse lookup tables to extract meeting room coordinates” (comprises “coordinates” (reference frame information) associated with each in-room user from the in-room device because according to ¶ 0028 line 21 “coordinates [are] in relation to the camera location”) “for each meeting participant”; ¶ 0057 S1: “At step 181, the method starts when a video/web conference call meeting is started” (the “video” (metadata) is captured at an interval which begins with the start of the conference) “in a conference room or area in which a video conference system is located”).
Regarding claim 5, Bhatt et al. do teach the device of claim 4, wherein the one or more time intervals comprises a start of a conferencing session (¶ 0057 S1: “At step 181, the method starts when a video/web conference call meeting is started” (the “video” (metadata) is captured at an interval which begins with the start of the conference) “in a conference room or area in which a video conference system is located”).
Regarding claim 6, Bhatt et al. do teach the device of claim 1, wherein the metadata information comprises a field of view (FOV) of the first camera and a resolution of the first camera (Abstract S1: “A video” (or “camera field of view” (metadata or camera information, ¶ 0035 2nd column line 9)) “conference call system is provided with a camera” (received in an in-room device) “to generate an input frame image of a conference room, where the video conference call system detects a human head”(detecting) “for each meeting participant” (one or more in-room users); ¶ 0035 S3: “And by knowing the camera field of view” (metadata) “resolution” (associated with the “camera” “resolution”) “in both horizontal and vertical directions with the respective horizontal and vertical pixel counts”).
Regarding claim 7, Bhatt et al. do teach the device of claim 1, wherein the processing circuitry is further configured to analyze a video stream coming from the in-room device (¶ 0059 S. before last: “In addition, one or more image” (video stream) “post-processing” (analysis) “steps may be applied to the results output from the head detection model”; note according to ¶ 0050 lines 4+: “head detector system 170” (using processing circuitry) “which processes” (analyzing) “incoming room-view video frame images” (video stream) “171 of a meeting room scene” (coming from the “camera” (the in-room device))).
Regarding claim 10, Bhatt et al. do teach a non-transitory computer-readable medium storing computer-executable instructions which when executed by one or more processors (¶ 0079 S2: “The disclosed system includes one or more first processors, a first data bus coupled to the one or more first processors, and a non-transitory, computer-readable storage medium embodying computer program code and being coupled to the first data bus, where the computer program code interacts with a plurality of computer operations and includes first instructions executable by the one or more first processors”)
Result in performing operations comprising:
identifying metadata comprising depth sensing information and camera information received from an in-room device located at a first location having a first camera (Abstract S1: “A video” (or “camera field of view” (metadata or camera information, ¶ 0035 2nd column line 9)) “conference call system is provided with a camera” (received in an in-room device) “to generate an input frame image of a conference room, where the video conference call system detects a human head”(detecting) “for each meeting participant” (one or more in-room users); ¶ 0002 S3: “detecting the location of meeting participant in a room location” “e.g., the distance” (including depth sensing information) “and direction between the camera and meeting participant” (associated with each in-room person); ¶ 0003 column 2 lines 7+: “identify” (identifying) “for each detected human head, a head bounding box with specified image plane coordinate and dimension information in a head box data structure” (a subset of the “video” (metadata)));
performing face recognition on one or more in-room users (Abstract lines 6-8: “generates a head bounding box which surround each detected human head” (for each face) “and identifies” (perform identification) “a corresponding meeting participant” (for each user of the one or more in-room users));
calculating a distance of a first in-room user based on the metadata and a first number of pixels across the face of the first in-room user (¶ 0048 S1: “determining the position and distance” (calculating a distance) “of a meeting participant” (of a first in-room user) “in a room based on the pixel count” (based on a first number of pixels) “for the height and width” (based on the “video” pictures (metadata)) “of each human head” (across his face));
calculating a distance between the first in-room user and a second in-room user based on the metadata and the first number of pixels across the face of the first in-room user and a second number of pixels across the face of the second in-room user (¶ 0078 first column last 3 lines+: “The disclosed methodology also applies the pixel width measure” (using the first number of pixels across each participant face) “and pixel height measure” “extracted from each head bounding box” (and metadata) “to one or more reverse lookup tables to extract meeting room coordinates” (determine distances of each in-room user from the in-room device because according to ¶ 0028 line 21 “coordinates [are] in relation to the camera location”) “for each meeting participant” (for e.g., a first in-room user as well as a second user, which enables calculating the distance between them, e.g., see Fig. 4 and ¶ 0031 lines 13+: “For example, a meeting participant 42” (e.g. a first in-room user) “located on the center line of sight of camera focal point” “at a distance of d=0.5 meters” “a meeting participant 43” (a second in-room user) “located on the center line of sight of camera” “at a distance of camera” “at a larger distance of d=1.0 (which readily gives the distance between “participant 42” (first in-room user) and “participant 43” (and second in-room user) as 0.5 meters)).
Bhatt et al. do not specifically disclose:
Generate a gain matrix for the one or more in-room users based, in part, on distances of the one or more in-room uses from the first camera and beamforming angles corresponding to the one or more in-room users; and
Apply gain values form the gain matrix to voice data corresponding to each of the one or more in-room users during post-processing on a downstream speaker playback path.
Herre et al. do teach:
Generate a gain matrix for the one or more in-room users based, in part, on distances of the one or more in-room uses from the first camera and beamforming angles corresponding to the one or more in-room users (page 27 lines 14+: “The distance-dependent gain” (gain depending on distance) “for each DoA is derived from the resulting length of the direction vector, dp(k,n)” (and direction vectors) and described as “Gi(rp(k,n))(||dp(k,n)||) -γ” (a gain matrix element dependent on “rp” (distance) and “dp” (direction), where “DoA” (“direction of arrival” (P.5 lines 3-4)) which “is derived” from “dp(k,n)” is according to page 9 lines 15-20 “calculate[ed]” “using” “meta data” which depends on “distance to a source” (distance from one or more users) “using two angles” (and beamforming angles) “with respect to two different reference locations and the distance positions”; page 4 lines 7-9: “arrays of cameras can be employed to generate light-field rendering” (e.g., using cameras to determine the distances and angles) “For audio, a similar set up employes distributed microphone arrays” (beamforming is used to help in audio “rendering”)) ; and
Apply gain values form the gain matrix to voice data corresponding to each of the one or more in-room users during post-processing on a downstream speaker playback path (page 3 lines 29-32: “The audio is often generated using object-based rendering” (during a post processing or playback of an “audio” (voice data)) “where each audio object is rendered with distance-dependent gain” (using the gain matrix gain values which depends on distance) “and relative direction” (and the beam forming angles) “from the user” (corresponding to one or more in room users) “based on the tracking data”).
It would have therefore been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to incorporate the “gain factor” incorporation into audio “rendering” (regeneration) of Herre et al. into the “sound processing” “to amplify the speech of [a] speaker who is located more than X meters form the camera” of Bhatt et al. (¶ 0027 page 3 lines 6-7) would enable the combined systems and their associated methods to perform in combination as they do separately and to further enable Bhatt et al. to achieve a “high-quality [audio] version” as disclosed in Herre et al. page 37 line 24 since Herre et al.’s “gain” depends on both “distance” and “direction”.
Regarding claim 11, Bhatt et al. do teach the non-transitory computer-readable medium of claim 10, wherein the depth sensing information is associated with the one or more in-room users located in a field of view of the first camera (¶ 0002 S3: “detecting the location of meeting participant in a room location” “e.g., the distance” (the depth sensing information) “and direction between the camera and meeting participant” (associated with each in-room user being located in the field of view of the “camera” (the first camera))).
Regarding claim 12, Bhatt et al. do teach the non-transitory computer-readable medium of claim 10, wherein the depth sensing information provides distances of the one or more in-room users (¶ 0002 S3: “detecting the location of meeting participant in a room location” “e.g., the distance” (the depth sensing information provides “distance” (distances)) “and direction between the camera and meeting participant” (for each in-room user from the in-room device)).
Regarding claim 13, Bhatt et al. do teach the non-transitory computer-readable medium of claim 10, wherein the metadata information comprises reference frame information associated with each of the one or more in-room users captured at one or more time intervals (¶ 0078 first column last 3 lines+: “The disclosed methodology also applies the pixel width measure” “and pixel height measure” “extracted from each head bounding box” (the metadata) “to one or more reverse lookup tables to extract meeting room coordinates” (comprises “coordinates” (reference frame information) associated with each in-room user from the in-room device because according to ¶ 0028 line 21 “coordinates [are] in relation to the camera location”) “for each meeting participant”; ¶ 0057 S1: “At step 181, the method starts when a video/web conference call meeting is started” (the “video” (metadata) is captured at an interval which begins with the start of the conference) “in a conference room or area in which a video conference system is located”).
Regarding claim 14, Bhatt et al. do teach the non-transitory computer-readable medium of claim 13, wherein the one or more time intervals comprises a start of a conferencing session (¶ 0057 S1: “At step 181, the method starts when a video/web conference call meeting is started” (the “video” (metadata) is captured at an interval which begins with the start of the conference) “in a conference room or area in which a video conference system is located”).
Regarding claim 15, Bhatt et al. do teach the non-transitory computer-readable medium of claim 10, wherein the metadata information comprises a field of view (FOV) of the first camera and a resolution of the first camera (Abstract S1: “A video” (or “camera field of view” (metadata or camera information, ¶ 0035 2nd column line 9)) “conference call system is provided with a camera” (received in an in-room device) “to generate an input frame image of a conference room, where the video conference call system detects a human head”(detecting) “for each meeting participant” (one or more in-room users); ¶ 0035 S3: “And by knowing the camera field of view” (metadata) “resolution” (associated with the “camera” “resolution”) “in both horizontal and vertical directions with the respective horizontal and vertical pixel counts”).
Regarding claim 16, Bhatt et al. do teach the non-transitory computer-readable medium of claim 10, wherein the operations further comprise analyze a video stream coming from the in-room device (¶ 0059 S. before last: “In addition, one or more image” (video stream) “post-processing” (analysis) “steps may be applied to the results output from the head detection model”; note according to ¶ 0050 lines 4+: “head detector system 170” (using processing circuitry) “which processes” (analyzing) “incoming room-view video frame images” (video stream) “171 of a meeting room scene” (coming from the “camera” (the in-room device))).
Regarding claim 19, Bhatt et al. do teach a method (Abstract S1: “A video conference call system” (a device and associated method) “is provided with a camera to generate an input frame image of a conference room”; ¶ 0081 lines 11+: “The information handling system may include random access memory (RAM), one or more processing resources such as a central processing unit (CPU) or hardware or software control logic”),
comprising:
identifying metadata comprising depth sensing information and camera information received from an in-room device located at a first location having a first camera (Abstract S1: “A video” (or “camera field of view” (metadata or first camera information, ¶ 0035 2nd column line 9)) “conference call system is provided with a camera” (received in an in-room device) “to generate an input frame image of a conference room, where the video conference call system detects a human head”(detecting) “for each meeting participant” (one or more in-room users); ¶ 0002 S3: “detecting the location of meeting participant in a room location” “e.g., the distance” (including depth sensing information) “and direction between the camera and meeting participant” (associated with each in-room person); ¶ 0003 column 2 lines 7+: “identify” (identifying) “for each detected human head, a head bounding box with specified image plane coordinate and dimension information in a head box data structure” (a subset of the “video” (metadata)));
performing face recognition on one or more in-room users (Abstract lines 6-8: “generates a head bounding box which surround each detected human head” (for each face) “and identifies” (perform identification) “a corresponding meeting participant” (for each user of the one or more in-room users));
calculating a distance of a first in-room user from the first camera based on the metadata and a first number of pixels across the face of the first in-room user (¶ 0048 S1: “determining the position and distance” (calculating a distance) “of a meeting participant” (of a first in-room user) “in a room based on the pixel count” (based on a first number of pixels) “for the height and width” (based on the “video” (obtained by the first camera) pictures (metadata)) “of each human head” (across his face));
calculating a distance between the first in-room user and a second in-room user based on the metadata and the first number of pixels across the face of the first in-room user and a number of pixels across the face of the second in-room user (¶ 0078 first column last 3 lines+: “The disclosed methodology also applies the pixel width measure” (using the first number of pixels across each participant face) “and pixel height measure” “extracted from each head bounding box” (and metadata) “to one or more reverse lookup tables to extract meeting room coordinates” (determine distances of each in-room user from the in-room device because according to ¶ 0028 line 21 “coordinates [are] in relation to the camera location”) “for each meeting participant” (for e.g., a first in-room user as well as a second user, which enables calculating the distance between them, e.g., see Fig. 4 and ¶ 0031 lines 13+: “For example, a meeting participant 42” (e.g. a first in-room user) “located on the center line of sight of camera focal point” “at a distance of d=0.5 meters” “a meeting participant 43” (a second in-room user) “located on the center line of sight of camera” “at a distance of camera” “at a larger distance of d=1.0 (which readily gives the distance between “participant 42” (first in-room user) and “participant 43” (and second in-room user) as 0.5 meters)).
Bhatt et al. do not specifically disclose:
Generate a gain matrix for the one or more in-room users based, in part, on distances of the one or more in-room uses from the first camera and beamforming angles corresponding to the one or more in-room users; and
Apply gain values form the gain matrix to voice data corresponding to each of the one or more in-room users during post-processing on a downstream speaker playback path.
Herre et al. do teach:
Generate a gain matrix for the one or more in-room users based, in part, on distances of the one or more in-room users from the first camera and beamforming angles corresponding to the one or more in-room users (page 27 lines 14+: “The distance-dependent gain” (gain depending on distance) “for each DoA is derived from the resulting length of the direction vector, dp(k,n)” (and direction vectors) and described as “Gi(rp(k,n))(||dp(k,n)||) -γ” (a gain matrix element dependent on “rp” (distance) and “dp” (direction), where “DoA” (“direction of arrival” (P.5 lines 3-4)) which “is derived” from “dp(k,n)” is according to page 9 lines 15-20 “calculate[ed]” “using” “meta data” which depends on “distance to a source” (distance from one or more users) “using two angles” (and beamforming angles) “with respect to two different reference locations and the distance positions”; page 4 lines 7-9: “arrays of cameras can be employed to generate light-field rendering” (e.g., using cameras to determine the distances and angles) “For audio, a similar set up employes distributed microphone arrays” (beamforming is used to help in audio “rendering”)) ; and
Apply gain values form the gain matrix to voice data corresponding to each of the one or more in-room users during post-processing on a downstream speaker playback path (page 3 lines 29-32: “The audio is often generated using object-based rendering” (during a post processing or playback of an “audio” (voice data)) “where each audio object is rendered with distance-dependent gain” (using the gain matrix gain values which depends on distance) “and relative direction” (and the beam forming angles) “from the user” (corresponding to one or more in room users) “based on the tracking data”).
It would have therefore been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to incorporate the “gain factor” incorporation into audio “rendering” (regeneration) of Herre et al. into the “sound processing” “to amplify the speech of [a] speaker who is located more than X meters form the camera” of Bhatt et al. (¶ 0027 page 3 lines 6-7) would enable the combined systems and their associated methods to perform in combination as they do separately and to further enable Bhatt et al. to achieve a “high-quality [audio] version” as disclosed in Herre et al. page 37 line 24 since Herre et al.’s “gain” depends on both “distance” and “direction”.
Regarding claim 20, Bhatt et al. do teach the method of claim 19, wherein the depth sensing information is associated with the one or more in-room users located in a field of view of the first camera (¶ 0002 S3: “detecting the location of meeting participant in a room location” “e.g., the distance” (the depth sensing information) “and direction between the camera and meeting participant” (associated with each in-room user being located in the field of view of the “camera” (the first camera))).
Regarding claim 21, Bhatt et al. do not specifically disclose the device of claim 1, wherein applying the gain values includes equalizing voice levels of the one or more in-room users.
Herre et al. do teach the device of claim 1, wherein applying the gain values includes equalizing voice levels of the one or more in-room users (Page 37 lines 20-21: “a monophonic” (equal in level) “sound signal” (audio of one or more in-room users) “is applied” (is obtained) “to a subset of loudspeakers after multiplication with loudspeaker-specific gain” (using the gain values) “factors”).
For obviousness to combine Bhatt et al. and Herre et al. see claim 1.
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claim(s) 8, 17 is/are rejected under 35 U.S.C. 103 as being unpatentable over Bhatt et al. in view of Herre et al., and further in view of Tangeland et al. (US 2023/0300295).
Regarding claim 8, Bhatt et al. in view of Herre et al. do not specifically disclose the device of claim 1, wherein the processing circuitry is further configured to:
select the first in-room user of the second in-room user by utilizing touch; and
Steer a beamformer in a direction of the first in-room user or the second in-room user.
Tangeland et al. do teach the device of claim 1 (Title Abstract)
Wherein the processing circuitry is further configured to:
Select the first in-room user or the second in-room user by utilizing touch (¶ 0028 S4: “In one embodiment, video endpoint device 120 may display the self-view of the conference room”(for in-room users) “on display 122 and, when display 122 is a touch screen, a participant may select” (select) “one of participants 202-210” (either a first or a second in-room user) “as the presenter by touching” (by utilizing touch) “an image of the presenter on the touch screen”);
And Steer a beamformer in a direction of the first in-room user or the second in-room user (¶ 0032 last S: “Microphones 320-1 to 320-N may detect audio from participants 302-310 and a position of a speaker” (towards a first and/or a second in-room user) “may be determined using the speaker tracking” (steering) “microphone array” (a beamformer); “tracking” a “speaker” requires “steering” a “microphone array” (“beamformer”) towards or in the direction of the “speaker”; i.e., see Ashoori et al. (US 2018/0315094) ¶ 0027 S1: “In some embodiments, multi-microphone arrays can dynamically steer “listening beams,” which, with the aid of video cameras, can track the location of the individual”).
It would have therefore been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to incorporate the methods of “speaker” “tracking” of Tangeland et al. in a “conference” to determination of “meeting room coordinates” for “each” “meeting participant” of Bhatt et al. in Bhatt et al. in view of Herre et al. would enable the combined systems and their associated methods to perform in combination as they do separately and to further enable Bhatt et al. in view of Herre et al. an alternative determination of “participant” “positions” to help assessing validity of its techniques.
Regarding claim 17, Bhatt et al. in view of Herre et al. do not specifically disclose the non-transitory computer-readable medium of claim 10, wherein the operations further comprise:
selecting the first in-room user or the second in-room user by utilizing touch; and
Steer a beamformer in a direction of the first in-room user or the second in-room user.
Tangeland et al. do teach:
Selecting the first in-room user or the second in-room user by utilizing touch (¶ 0028 S4: “In one embodiment, video endpoint device 120 may display the self-view of the conference room”(for in-room users) “on display 122 and, when display 122 is a touch screen, a participant may select” (select) “one of participants 202-210” (either a first or a second in-room user) “as the presenter by touching” (by utilizing touch) “an image of the presenter on the touch screen”);
And Steer a beamformer in a direction of the first in-room user or the second in-room user (¶ 0032 last S: “Microphones 320-1 to 320-N may detect audio from participants 302-310 and a position of a speaker” (towards a first and/or a second in-room user) “may be determined using the speaker tracking” (steering) “microphone array” (a beamformer); “tracking” a “speaker” requires “steering” a “microphone array” (“beamformer”) towards or in the direction of the “speaker”; i.e., see Ashoori et al. (US 2018/0315094) ¶ 0027 S1: “In some embodiments, multi-microphone arrays can dynamically steer “listening beams,” which, with the aid of video cameras, can track the location of the individual”).
It would have therefore been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to incorporate the methods of “speaker” “tracking” of Tangeland et al. in a “conference” to determination of “meeting room coordinates” for “each” “meeting participant” of Bhatt et al. in Bhatt et al. in view of Herre et al. would enable the combined systems and their associated methods to perform in combination as they do separately and to further enable Bhatt et a. in view of Herre et al. an alternative determination of “participant” “positions” to help assessing validity of its techniques.
Claim(s) 9, 18 is/are rejected under 35 U.S.C. 103 as being unpatentable over Bhatt et al. in view of Herre et al., and further in view of Foote et al. (US 2005/0007445).
Regarding claim 9, Bhatt et al. in view of Herre et al. do not specifically disclose the device of claim 1, wherein the processing circuitry is further configured to:
monitor at least one of the one or more in-room users using gaze; and
enhance voice data of the at least one of the one or more in-room users by direction-based tuning.
Foote et al. do teach the device of claim 1 (Title, Abstract),
wherein the processing circuitry is further configured to:
monitor at least one of the one or more in-room users using gaze (Abstract last S: “A camera can be mounted adjacent to the screen, and can allow the subject to view a selected conference participant or a desired location such that when the camera is trained on the selected participant” (monitor at least one or more of “conference participant[s]” (in-room users)) “or desired location a gaze” (using gaze) “of the remote participant displayed by the screen appears substantially directed at the selected participant or desired location”); and
enhance voice data of the at least one of the one or more in-room users by direction-based tuning (¶ 0033 S3: “A microphone array 114 can provide directional speech pickup” (a direction-based voice tuning) “and enhancement over a range of participant” (on e.g. at least one or more in-room users) “positions”).
It would have therefore been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to incorporate the overall “VIDEO TELECONFERENCING” methods of Foote et al. into the “video conference call system” of Bhatt et al. in Bhatt et al. in view of Herre et al. , 4-7, would enable the combined systems and their associated methods to perform in combination as they do separately and to further enable Bhatt et al. to “provide directional speech” “enhancement over a range of [a] participant positions” as disclosed in Foote et al. ¶ 0033 S3 to help Bhatt et al. “sound processing” “to amplify the speech of a speaker” as disclosed in Bhatt et al. ¶ 0027 2nd column lines 6-7.
Regarding claim 18, Bhatt et al. in view of Herre et al. do not specifically disclose the non-transitory computer-readable medium of claim 10, wherein the operations further comprise:
monitor at least one of the one or more in-room users using gaze; and
enhance voice data of the at least one of the one or more in-room users by direction-based tuning.
Foote et al. do teach:
monitor at least one of the one or more in-room users using gaze (Abstract last S: “A camera can be mounted adjacent to the screen, and can allow the subject to view a selected conference participant or a desired location such that when the camera is trained on the selected participant” (monitor at least one or more of “conference participant[s]” (in-room users)) “or desired location a gaze” (using gaze) “of the remote participant displayed by the screen appears substantially directed at the selected participant or desired location”); and
enhance voice data of the at least one of the one or more in-room users by direction-based tuning (¶ 0033 S3: “A microphone array 114 can provide directional speech pickup” (a direction-based voice tuning) “and enhancement over a range of participant” (on e.g. at least one or more in-room users) “positions”).
It would have therefore been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to incorporate the overall “VIDEO TELECONFERENCING” methods of Foote et al. into the “video conference call system” of Bhatt et al. in Bhatt et al. in view of Herre et al. would enable the combined systems and their associated methods to perform in combination as they do separately and to further enable Bhatt et al. in view of Herre et al. to “provide directional speech” “enhancement over a range of [a] participant positions” as disclosed in Foote et al. ¶ 0033 S3 to help Bhatt et al. “sound processing” “to amplify the speech of a speaker” as disclosed in Bhatt et al. ¶ 0027 2nd column lines 6-7.
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure. Xu et al. (US Patent 10,943,204 Col. 5 lines 18+).
Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a).
A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to FARZAD KAZEMINEZHAD whose telephone number is (571)270-5860. The examiner can normally be reached 10:30 am to 11:30 pm.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Paras D. Shah can be reached at (571) 270-1650. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/Farzad Kazeminezhad/
Art Unit 2653
July 16th 2026.