Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claims 1, 14, 24, 28-29 are rejected under 35 U.S.C. 103 as being unpatentable over OZ (US 20210360195) in view of Haaland (US 20210274129).
Regarding claim 1, OZ teaches a computer implemented method of virtualization of one or more sensors (Paragraph 58: representation of a virtual 3D video conference environment that is associated with the participant is a representation that is shown to the participant. Different participants may be associated with different representation of a virtual 3D video conference environment), comprising:
defining a location of a virtual sensor “in paragraph 124: one or more sensors such as camera” and (Paragraph 274: correction may correct any deviations between the actual optical axis of the camera and a desired optical axis of a virtual camera. While some of the example refer to the height of the virtual camera, any of the following may also refer to the lateral location of the camera—for example positioning the virtual camera at the center of the display (both height and lateral location, positioning the virtual camera to have a virtual optical axis that is directed to the eyes of a participant (for example via a virtual optical axis that may be perpendicular to the display or have any other spatial relationship with the display);
analyzing a plurality of sensor datasets of the sensor array captured from a plurality of different views (Paragraph 168: Generative Adversarial Network (GAN) may be trained based on many images of a certain person or on many images of multiple people to generate images of people from angles that may be different than the angle at which the camera may be currently seeing the person);
computing a virtual sensor dataset for the location of the virtual sensor according to the analysis (Paragraph 169: At runtime, such a network would receive an image of a person as an input and a camera position from which the person should be rendered. The network would render an image of that person from the different camera position including parts that may be obscured in the input image or may be at a low resolution in the input image due to being almost parallel to the camera's line of sight (i.e. cheeks at a frontal image);
and directing a stream of the virtual sensor dataset to a client terminal (Paragraph 64: step 250 of transmitting the updated representation of virtual 3D video conference environment to at least one device of at least one participant).
Oz does not explicitly teach defining a location of a virtual sensor on a sensor array disposed on a common surface.
Haaland in the same art of endeavor teaches defining a location of a virtual sensor on a sensor array disposed on a common surface (Paragraph 6: capturing data of a first party at a first location using an array of one or more video cameras and/or one or more sensors; [0007] determining, for each of the one or more video cameras and/or each of the one or more sensors in the array, the three-dimensional position(s) of one or more features represented in the data captured by the video camera or sensor; [0008] defining a virtual camera positioned at a three-dimensional virtual camera position; [0009] transforming the three-dimensional position(s) determined for the feature(s) represented in the data into a common coordinate system to form a single view of the feature(s) as appearing to have been captured from the virtual camera using the video image data from the one or more video cameras and/or the data captured by the one or more sensors ).
Therefore, it would have been obvious to one with ordinary skill in the art to modify OZ with Haaland in order to improve the system and determine the direction of arrival of a signal, increase Redundancy and Reliability and achieve Higher Signal-to-Noise Ratio (SNR) and Array Gain.
Regarding claim 14, OZ in view of Haaland teaches wherein analyzing comprises feeding the plurality of sensor datasets into a virtualization machine learning model that generates the virtual sensor dataset as an outcome thereof (OZ: Paragraph 342).
Regarding claim 24, OZ in view of Haaland teaches wherein the virtual sensor dataset includes a depth map indicating a respective depth for each data element of the virtual sensor dataset relative to a normal at a location of the virtual sensor (OZ: Paragraph 165: If the camera is a 3D depth camera, then the depth data can be used to make the models more accurate and solve ambiguities. For example, if one obtains only a front facing image of a person's head, it may be impossible to know the exact depth of each point in the image, i.e. the length of the nose. When more than one image of the face from different angles exist, then such ambiguities may be solved. Nevertheless, there may remain occluded areas seen in only one image or inaccuracies. The depth data from a depth camera may assist in generating a 3D model with depth information in every point that solves the ambiguity problems, Paragraph 209] b. A high frequency depth map detailing such fine details as wrinkles, skin moles, etc. [0210] c. A reflectance map detailing the color of each part of the face or body. Multiple reflectance maps may be used to model the change of appearance from different angles. [0211] d. An optional material map detailing the material from which each polygon may be made, e.g. skin, hair, cloth, plastic, metal, etc. [0212] e. An optional semantic map listing what part of the body each part in the 3D model or reflectance map represents. [0213] f. These models and maps may be created before the meeting, during the meeting or may be a combination or models created before and during the meeting).
Regarding claim 28, OZ in view of Haaland teaches, wherein the sensor array comprises a plurality of sensors selected from a group comprising imaging sensors and audio sensors (OZ: abstract).
Regarding claim 29, see claim 1 rejection.
Allowable Subject Matter
Claims 2-6, 8, 15-16, 21-23, 25-27 are objected to as being dependent upon a rejected base claim, but would be allowable if rewritten in independent form including all of the limitations of the base claim and any intervening claims.
During search examiner found the following arts related to claim 2:
OZ teaches detecting the movement of the face, expression and change location of the virtual camera (142] A 3D model composed of a fixed number of triangles and vertices may be deformed as the 3D model changes. For example, a 3D model of a face may be deformed as the face changes its expression. Nevertheless, the pixels in the texture map correspond to the same locations in the same triangles, even though the 3D locations of the triangles change as the expression of the face changes.
Yerli (US 20230316663) teaches associating the view of the first virtual camera to the coordinates of the key facial landmarks; tracking movement of the key facial landmarks in 6 degrees of freedom based on the movement of the head of the first user; and adjusting the position and orientation of the first virtual camera based on the tracked movement of the key facial landmarks. The method may further include dynamically selecting the elements of the virtual environment based on the adjusted position of the virtual camera; and presenting the selected elements of the virtual environment to the corresponding client device. (Paragraph 8 and 15).
Wee (US 20050080849) teaches detect the movement of the individual and select new imaging device to capture the images (Paragraph 32: The communication provider 18 may employ machine vision or audio processing techniques to detect movements of the individual involved in an interest thread. In response, the communication provider 18 selects a new set of sensing and rendering components for the interest thread based on the new locations of the individuals involved in the interest thread and the specified coverage areas of the available sensing and rendering components).
None of the cited references alone or in reasonable combination teaches this (dynamically tracking motion of the user across the display to identify a new location of the user on the display; dynamically selecting a new location of the virtual sensor on the display according to the new location of the user on the display; computing the virtual sensor dataset for the new location of the virtual sensor; and directing the stream of the virtual sensor dataset to the client terminal) as claimed in claim 2.
Claims 3-6 are objected as depending on claim 2.
Also none of the found, cited arts alone or in combination teaches (virtual sensor is selected for generating a single face-on view of a first participant as seen remotely by at least one second participant participating in a conference, wherein the plurality of sensor datasets of the sensor array capture the first participant from the plurality of different views, and wherein the client terminal is of the at least one second participant, for depicting a single face-on view of the first participant as viewed from the virtual sensor by the second participant) as claimed by claim 8.
Claim 9 depends on claim 8.
Also, none of the cited arts alone or in reasonable combination teaches (wherein the generator component generates an outcome of virtual sensor dataset corresponding to a virtual sensor in response to an input of a plurality of sensor datasets captured by a plurality of sensors, wherein the generating component is adapted during training on feedback from the discriminator component for generating virtual sensor datasets that the discriminator component cannot accurately distinguish from ground truth datasets captured by a ground truth sensor) as claimed by claim 15.
Claim 16 depend on claim 15.
Also, none of the cited arts alone or in reasonable combination teaches (each sensor of the sensor array is configured for outputting a compromised quality dataset depicting a partial field of view, and computing the virtual sensor dataset comprises stitching a plurality of compromised quality datasets into a main dataset at higher quality that depicts a full field of view) as claimed in claim 21.
Claims 22-23 depends on claim 21.
Also, none of the cited arts alone or in reasonable combination Teaches (the depth map is computed by feeding the plurality of sensor datasets into a depth ML model trained on a depth training dataset comprising a plurality of records, each record including a plurality of sample sensor datasets and a ground truth of a sample depth map) as claimed in claim 25.
Claim 26 depend on claim 25.
Also, none of the cited arts alone or in reasonable combination Teaches (merging the plurality of sensor datasets into a lattice in which each sensor dataset is positioned based on location in space of the corresponding sensor, wherein the lattice represents a single dataset in which relative location of each sensor is defined; and converting a plurality of positions of the lattice into indications of depth relative to the virtual sensor for computing the depth map for the virtual sensor dataset).
Conclusion
Any inquiry concerning this communication or earlier communications from the examiner should be directed to MARIA EL-ZOOBI whose telephone number is (571)270-3434. The examiner can normally be reached Monday-Friday 7-4.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Carolyn Edward can be reached at (571)270-7136. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/MARIA EL-ZOOBI/Primary Examiner, Art Unit 2692