Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Information Disclosure Statement
The information disclosure statement (IDS) submitted on 06/16/2025 was filed after the mailing date of the non-final rejection on 07/15/2026. The submission is in compliance with the provisions of 37 CFR 1.97. Accordingly, the information disclosure statement is being considered by the examiner.
Claim Rejections - 35 USC § 101
35 U.S.C. 101 reads as follows:
Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title.
Claim 25 is rejected under 35 U.S.C. 101 as not being patentable material.
Regarding Claim 25, it falls under the exclusions to patentability listed in MPEP 2106.03(I). Specifically, "Products that do not have a physical or tangible form, such as information (often referred to as 'data per se') or a computer program per se (often referred to as 'software per se') when claimed as a product without any structural recitations.” To remedy this, Examiner recommends amending the claim language to open with “A non-transitory, computer-readable storage medium containing a program that, when executed [insert function here].”
Claim Rejections - 35 USC § 102
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action:
A person shall be entitled to a patent unless –
(a)(1) the claimed invention was patented, described in a printed publication, or in public use, on sale, or otherwise available to the public before the effective filing date of the claimed invention.
(a)(2) the claimed invention was described in a patent issued under section 151, or in an application for patent published or deemed published under section 122(b), in which the patent or application, as the case may be, names another inventor and was effectively filed before the effective filing date of the claimed invention.
Claims 1, 13, & 25 are rejected under 35 U.S.C. 102(a)(1) as being anticipated by Sheng & Liu (CN 113542604 A, hereinafter, "Sheng").
Regarding Claim 1, Sheng teaches a method for focusing a camera capturing video data, the method being performed in a focus determiner (Sheng, Abstract, ln. 1, "…a video focusing method…"), the method comprising: obtaining audio data and video data, comprising a plurality of sequential images, wherein the audio data and video data overlap in time and cover the same scene (Sheng, Abstract, ln. 4-5, "…identifying the target object matched with the target voice print characteristic in the video scene, focusing or tracking the face of the target object."); determining directional data of a dominant sound source in the audio data (Sheng, Abstract, ln. 4-5, "…identifying the target object matched with the target voice print characteristic in the video scene, focusing or tracking the face of the target object."); matching a visual feature spatially, in an image of the plurality of sequential images of the video data, with the directional data, wherein the image corresponds to the directional data in time (Sheng, Abstract, ln. 4-5, "…identifying the target object matched with the target voice print characteristic in the video scene, focusing or tracking the face of the target object."); and focusing a camera, being the source of the video data, to focus on the visual feature for subsequent capturing of video data (Sheng, Abstract, ln. 4-5, "…identifying the target object matched with the target voice print characteristic in the video scene, focusing or tracking the face of the target object.").
Regarding Claim 13, Sheng teaches A focus determiner for focusing a camera capturing video data (Sheng, Abstract, ln. 1, "…a video focusing method…"), the focus determiner comprising: a processor (Sheng, pg. 4, para. 3, ln. 1-2, "…a processor…"); and a memory storing instructions that, when executed by the processor, cause the focus determiner to: obtain audio data and video data, comprising a plurality of sequential images, wherein the audio data and video data overlap in time and cover the same scene (Sheng, Abstract, ln. 4-5, "…identifying the target object matched with the target voice print characteristic in the video scene, focusing or tracking the face of the target object."); determine directional data of a dominant sound source in the audio data (Sheng, Abstract, ln. 4-5, "…identifying the target object matched with the target voice print characteristic in the video scene, focusing or tracking the face of the target object."); match a visual feature spatially, in an image of the plurality of sequential images of the video data, with the directional data, wherein the image corresponds to the directional data in time (Sheng, Abstract, ln. 4-5, "…identifying the target object matched with the target voice print characteristic in the video scene, focusing or tracking the face of the target object."); and focus a camera, being the source of the video data, to focus on the visual feature for subsequent capturing of video data (Sheng, Abstract, ln. 4-5, "…identifying the target object matched with the target voice print characteristic in the video scene, focusing or tracking the face of the target object.").
Regarding Claim 25, Sheng teaches a computer program for focusing a camera capturing video data, the computer program comprising computer program code which, when executed on a focus determiner causes the focus determiner to: obtain audio data and video data, comprising a plurality of sequential images, wherein the audio data and video data overlap in time and cover the same scene (Sheng, Abstract, ln. 4-5, "…identifying the target object matched with the target voice print characteristic in the video scene, focusing or tracking the face of the target object."); determine directional data of a dominant sound source in the audio data (Sheng, Abstract, ln. 4-5, "…identifying the target object matched with the target voice print characteristic in the video scene, focusing or tracking the face of the target object."); match a visual feature spatially, in an image of the plurality of sequential images of the video data, with the directional data, wherein the image corresponds to the directional data in time (Sheng, Abstract, ln. 4-5, "…identifying the target object matched with the target voice print characteristic in the video scene, focusing or tracking the face of the target object."); and focus a camera, being the source of the video data, to focus on the visual feature for subsequent capturing of video data (Sheng, Abstract, ln. 4-5, "…identifying the target object matched with the target voice print characteristic in the video scene, focusing or tracking the face of the target object.").
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claims 2, 8, 14, 20, & 24 are rejected under 35 U.S.C. 103 as being unpatentable over Sheng in view of Ugur & Tammi (US 20180084365 A1, hereinafter, "Ugur").
Regarding Claim 2, Sheng teaches the limitations of dependent Claim 1 as noted above. Ugur teaches the determining directional data comprises determining a direction of the dominant sound source based on multiple sound channels in the audio data (Ugur, [0080], ln. 1-3, "Multichannel systems, such as commonly used 5.1 channel setup, can be used for representing spatial signals with sound sources in different directions and thus they can potentially be used for representing the spatial events captured by a multi-microphone system."). It would have been obvious to a person having ordinary skill in the art at the time of the invention to combine the teachings of Ugur with those of Sheng because it is widely known in the art to use multi-microphone systems to triangulate the direction of an audio source.
Regarding Claim 8, Sheng teaches the limitations of dependent Claim 1 as noted above. Ugur teaches the camera is provided in a user device (Ugur, Fig. 1, 10), and wherein the method further comprises: performing voice identification on the audio data, wherein the voice identification results in a match with a person (Ugur, Fig. 8, [0266], ln. 2-3, "…the image 1403 comprises a first audio source, a speaker or person speaker shown by the highlighted selection box 1401."); and wherein the matching a visual features comprises increasing a matching priority for a sound source matching directional data of the person (Ugur, [0276], ln. 1-2, "…the user can further provide a user input indicating whether the object should be amplified or attenuated by pressing a corresponding icon on the screen."), when the person is associated with the user of the user device, but fails to be the user (Ugur, [0274], ln. 1-3, "…the user can in some embodiments then activate an object selection by pressing a dedicated icon on the screen and indicating an interesting object by selecting ‘tapping’ it."). It would have been obvious to a person having ordinary skill in the art at the time of the invention to combine the teachings of Ugur with those of Sheng because it is widely known in the art to include a camera in a user device, match a voice to a person, and use the sound source and directional data of said person to match visual features when the person is not the user.
Regarding Claim 14, Sheng teaches the limitations of dependent Claim 13 as noted above. Ugur teaches the instructions to determine directional data comprise instructions that, when executed by the processor, cause the focus determiner to determine a direction of the dominant sound source based on multiple sound channels in the audio data (Ugur, [0080], ln. 1-3, "Multichannel systems, such as commonly used 5.1 channel setup, can be used for representing spatial signals with sound sources in different directions and thus they can potentially be used for representing the spatial events captured by a multi-microphone system.").
Regarding Claim 20, Sheng teaches the limitations of dependent Claim 13 as noted above. Ugur teaches the camera is provided in a user device, and wherein the focus determiner further comprises instructions that, when executed by the processor, cause the focus determiner to perform voice identification on the audio data, wherein the voice identification results in a match with a person (Ugur, Fig. 8, [0266], ln. 2-3, "…the image 1403 comprises a first audio source, a speaker or person speaker shown by the highlighted selection box 1401."); and wherein the instructions to match a visual features comprise instructions that, when executed by the processor, cause the focus determiner to increasing a matching priority for a sound source matching directional data of the person (Ugur, [0276], ln. 1-2, "…the user can further provide a user input indicating whether the object should be amplified or attenuated by pressing a corresponding icon on the screen."), when the person is associated with the user of the user device, but fails to be the user (Ugur, [0274], ln. 1-3, "…the user can in some embodiments then activate an object selection by pressing a dedicated icon on the screen and indicating an interesting object by selecting ‘tapping’ it.").
Regarding Claim 24, Sheng teaches the limitations of dependent Claim 13 as noted above. Ugur teaches the camera and at least one microphone, for capturing the audio data, are mounted in a fixed relation to each other (Ugur, [0080], ln. 1-3, "Multichannel systems, such as commonly used 5.1 channel setup, can be used for representing spatial signals with sound sources in different directions and thus they can potentially be used for representing the spatial events captured by a multi-microphone system.").
Claims 3, 4, 11, 15, & 16 are rejected under 35 U.S.C. 103 as being unpatentable over Sheng in view of Lee (US 20140354874 A1, hereinafter, "Lee").
Regarding Claim 3, Sheng teaches the limitations of dependent Claim 1 as noted above. Lee teaches the matching a visual feature comprises classifying a plurality of objects in the image being potential sound sources, and determining the visual feature to be the one of the plurality of objects that is located in a position that best matches the directional data (Lee, [0059], ln. 13-15, "An algorithm for measuring a distance to the subject using the symbol recognition sensor may measure a direction of the subject through a microphone sensor first, may set the corresponding direction to be an ROI…"). It would have been obvious to a person having ordinary skill in the art at the time of the invention to combine the teachings of Lee with those of Sheng because it is widely known in the art to use microphones to determine the direction of a person and use that direction to set a corresponding Region of Interest (ROI).
Regarding Claim 4, Sheng and Lee teach the limitations of dependent Claim 3 as noted above. Sheng teaches the plurality of objects are faces (Sheng, [0163], ln. 12-13, "For example the image may detect a shape or colour which the apparatus is to track—for example a face."). It would have been obvious to a person having ordinary skill in the art at the time of the invention to combine the teachings of Sheng with those of Sheng and Lee because it is widely known in the art for faces to be an ROI.
Regarding Claim 11, Sheng teaches the limitations of dependent Claim 1 as noted above. Lee teaches performing speech recognition of at least one voice sound source in the audio data, resulting in text data for each one of the at least one voice sound source (Lee, [0032], ln. 2-3, "The wireless signal may include data provided in various forms as a voice call signal, a video call signal, and a text/multimedia message is transmitted and received."). It would have been obvious to a person having ordinary skill in the art at the time of the invention to combine the teachings of Lee with those of Sheng because it is widely known in the art to utilize speech recognition to convert speech to text data. Lee does not teach the matching a visual feature comprises adjusting a matching priority for each one of the at least one voice sound source based on the respective text data. However, Sheng teaches the matching a visual feature comprises adjusting a matching priority for each one of the at least one voice sound source based on the respective text data (Sheng, Abstract, ln. 4-5, "…identifying the target object matched with the target voice print characteristic in the video scene, focusing or tracking the face of the target object." The text data is based on the audio data. Therefore, the matching priority is based on the audio data as well.). It would have been obvious to a person having ordinary skill in the art at the time of the invention to combine the teachings of Sheng with those of Sheng and Lee because it is widely known in the art to identify persons matched with voice characteristics for utilizing speak-to-text functions.
Regarding Claim 15, Sheng teaches the limitations of dependent Claim 13 as noted above. Lee teaches the instructions to match a visual feature comprise instructions that, when executed by the processor, cause the focus determiner to classify a plurality of objects in the image being potential sound sources, and determine the visual feature to be the one of the plurality of objects that is located in a position that best matches the directional data (Lee, [0059], ln. 13-15, "An algorithm for measuring a distance to the subject using the symbol recognition sensor may measure a direction of the subject through a microphone sensor first, may set the corresponding direction to be an ROI…").
Regarding Claim 16, Sheng and Lee teach the limitations of dependent Claim 15 as noted above. Sheng teaches the plurality of objects are faces (Sheng, [0163], ln. 12-13, "For example the image may detect a shape or colour which the apparatus is to track—for example a face.").
Claims 5 & 17 are rejected as being unpatentable over Sheng in view of Lee and Feng et al (US 20110285808 A1, hereinafter, "Feng").
Regarding Claim 5, Sheng and Lee teach the limitations of dependent Claim 4 as noted above. Feng teaches determining that a conversation occurs between a plurality of people associated with the faces, and wherein the matching a visual feature comprises increasing a matching priority for sound sources matching directional data of the plurality of people associated with the faces (Feng, [0104], ln. 4-7, "…speaker recognition techniques and historical information can be used to identify the speaker based on their speech characteristics. Then, the endpoint 10 can steer the camera 50B to the last location associated with that recognized speaker, as long as it at least matches the speaker's current location."). It would have been obvious to a person having ordinary skill in the art at the time of the invention to combine the teachings of Feng with those of Sheng and Lee because it is widely known in the art to match a person to speech characteristics and match said person and their location to said speech characteristics.
Regarding Claim 17, Sheng and Lee teach the limitations of dependent Claim 16 as noted above. Feng teaches determine that a conversation occurs between a plurality of people associated with the faces, and wherein the instructions to match a visual feature comprise instructions that, when executed by the processor, cause the focus determiner to increase a matching priority for sound sources matching directional data of the plurality of people associated with the faces (Feng, [0104], ln. 4-7, "…speaker recognition techniques and historical information can be used to identify the speaker based on their speech characteristics. Then, the endpoint 10 can steer the camera 50B to the last location associated with that recognized speaker, as long as it at least matches the speaker's current location.").
Claims 6, 7, 18, & 19 are rejected under 35 U.S.C. 103 as being unpatentable over Sheng in view of Feng.
Regarding Claim 6, Sheng teaches the limitations of dependent Claim 1 as noted above. Feng teaches adjusting a field of view such that the video data covers the dominant directional data of the dominant sound source (Feng, Fig. 5, [0105], ln. 1-3, "Once the speaker is located, the endpoint 10 converts the speaker's candidate location into camera commands [pan-tilt-zoom coordinates] to steer the people-view camera 50B to capture the speaking participant [Block 214]."). It would have been obvious to a person having ordinary skill in the art at the time of the invention to combine the teachings of Feng with those of Sheng because it is widely known in the art to locate a person who is speaking and steer the camera to capture said person.
Regarding Claim 7, Sheng teaches the limitations of dependent Claim 1 as noted above. Feng teaches the method is repeated periodically with a time period corresponding to a preconfigured number of frames of the video data (Feng, [0155], ln. 6-7, "For example, the frame rate may be decimated to about six frames per second in one implementation."). It would have been obvious to a person having ordinary skill in the art at the time of the invention to combine the teachings of Feng with those of Sheng because it is widely known in the art to perform video functions at set intervals.
Regarding Claim 18, Sheng teaches the limitations of dependent Claim 13 as noted above. Feng teaches adjust a field of view such that the video data covers the dominant directional data of the dominant sound source (Feng, Fig. 5, [0105], ln. 1-3, "Once the speaker is located, the endpoint 10 converts the speaker's candidate location into camera commands [pan-tilt-zoom coordinates] to steer the people-view camera 50B to capture the speaking participant [Block 214].").
Regarding Claim 19, Sheng teaches the limitations of dependent Claim 13 as noted above. Feng teaches the instructions are repeated periodically with a time period corresponding to a preconfigured number of frames of the video data (Feng, [0155], ln. 6-7, "For example, the frame rate may be decimated to about six frames per second in one implementation.").
Allowable Subject Matter
Claim 9 is objected to as being dependent upon a rejected base claim, but would be allowable if rewritten in independent form including all of the limitations of the base claim and any intervening claims.
Regarding Claim 9, the prior art of record – taken alone or in combination – fails to teach or render obvious the determining directional data of a dominant sound source comprises excluding, as a potential dominant sound source, a sound source matching directional data of the person, when the voice data of the person is a matched with voice data of the user of the user device.
Conclusion
Any inquiry concerning this communication or earlier communications from the examiner should be directed to STEVEN DANIEL BARRY whose telephone number is (571)270-0432. The examiner can normally be reached M-Th 0730-1630.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Lin Ye can be reached on 517-272-7372. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/STEVEN DANIEL BARRY/Examiner, Art Unit 2638
/LIN YE/Supervisory Patent Examiner, Art Unit 2638