DETAILED ACTION
This action is in response to the application filed 01/17/2025. Claims 1 – 5 are pending and have
been examined.
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Claim Objections
Claim 3 is objected to because of the following informalities:
Claim 3 states “in the fixed mode, the processing circuitry is further configured to execute the 3D display process such that an apparent position of the first user and and an apparent position of the second user are present on lines connecting between the left and right eyes of the third user and the display surface”. This appears to be a typographical error and should state “in the fixed mode, the processing circuitry is further configured to execute the 3D display process such that an apparent position of the first user and an apparent position of the second user are present on lines connecting between the left and right eyes of the third user and the display surface”.
Appropriate correction is required.
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claims 1 and 5 are rejected under 35 U.S.C. 103 as being unpatentable over Azuma et al. (U.S. Pub. No. 2024/0015263, hereinafter “Azuma”) in view of Oz et al. (U.S. Pub. No. 2023/0171380, hereinafter “Oz”).
Regarding Claim 1, Azuma teaches
An electronic device for displaying information (see Azuma Paragraph [0151], FIG. 16 is a block diagram of an example processor platform 1600 structured to execute and/or instantiate the machine readable instructions and/or operations of FIGS. 11, 12, 13, 14, and 15 to implement the example telepresence circuitry 408 of FIG. 4. The processor platform 1600 can be, for example, a server, a personal computer, a workstation, a self-learning machine (e.g., a neural network), a mobile device (e.g., a cell phone, a smart phone, a tablet such as an iPad™), a personal digital assistant (PDA), an Internet appliance, a DVD player, a CD player, a digital video recorder, a Blu-ray player, a gaming console, a personal video recorder, a set top box, a headset (e.g., an augmented reality (AR) headset, a virtual reality (VR) headset, etc.) or other wearable device, or any other type of computing device), comprising:
a display having a display surface (see Azuma Paragraph [0047],FIG. 4 is a block diagram of an example system to provide augmented telepresence communication. The example system 400 includes users 402A, 402B, 402C, and 402D, example camera arrays 404A, 404B, 404C, and 404D, network 406, example telepresence circuitry 408, displays 410A, 410B, 410C, and 410D, and microphones 412A, 412B, 412C, and 412D); and
processing circuitry (see Azuma Paragraph [0056], FIG. 5 is a block diagram of an example implementation of example telepresence circuitry to provide remote telepresence communication. The example telepresence circuitry 408 of FIG. 5 may be instantiated (e.g., creating an instance of, bring into being for any length of time, materialize, implement, etc.) by processor circuitry such as a central processing unit executing instructions), wherein
a first user image is an image representing a first user who is a user of a first electronic device communicating with the electronic device (see Azuma Paragraph [0049], The example camera arrays 404 capture a subset of the example user's light field from a plurality of views. The subset of the example user's light field is captured as a plurality of images. The plurality of images represent individual frames from a plurality of videos, where each of the cameras that compose a given camera array 404A records a video. The example camera arrays 404 provide the plurality of images to the example telepresence circuitry 408 via the network 406. The users 402 in the example system 400 utilize respective example camera arrays 404. For example, camera array 404A provides images of user 402A, camera array 404B provides images of user 402B, etc., Paragraph [0175], identify features from a plurality of images, the plurality of images representing a first user and a second user, social view synthesis circuitry to create a first representation of the first user and a second representation of the second user, the representations created using the plurality of images, the representations representing the respective users at specified distances and specified perspectives from a viewer, Paragraph [0182], the plurality of images represent video feeds from a first camera array used by the first user and a second camera array used by the second user),
a second user image is an image representing a second user who is a user of a second electronic device communicating with the electronic device (see Azuma Paragraph [0049], The example camera arrays 404 capture a subset of the example user's light field from a plurality of views. The subset of the example user's light field is captured as a plurality of images. The plurality of images represent individual frames from a plurality of videos, where each of the cameras that compose a given camera array 404A records a video. The example camera arrays 404 provide the plurality of images to the example telepresence circuitry 408 via the network 406. The users 402 in the example system 400 utilize respective example camera arrays 404. For example, camera array 404A provides images of user 402A, camera array 404B provides images of user 402B, etc., Paragraph [0175], identify features from a plurality of images, the plurality of images representing a first user and a second user, social view synthesis circuitry to create a first representation of the first user and a second representation of the second user, the representations created using the plurality of images, the representations representing the respective users at specified distances and specified perspectives from a viewer, Paragraph [0182], the plurality of images represent video feeds from a first camera array used by the first user and a second camera array used by the second user), and
the processing circuitry is configured to (see Azuma Paragraph [0056], FIG. 5 is a block diagram of an example implementation of example telepresence circuitry to provide remote telepresence communication. The example telepresence circuitry 408 of FIG. 5 may be instantiated (e.g., creating an instance of, bring into being for any length of time, materialize, implement, etc.) by processor circuitry such as a central processing unit executing instructions. Additionally or alternatively, the example telepresence circuitry 408 of FIG. 5 may be instantiated (e.g., creating an instance of, bring into being for any length of time, materialize, implement, etc.) by an ASIC or an FPGA structured to perform operations corresponding to the instructions. It should be understood that some or all of the circuitry of FIG. 5 may, thus, be instantiated at the same or different times. Some or all of the circuitry may be instantiated, for example, in one or more threads executing concurrently on hardware and/or in series on hardware. Moreover, in some examples, some or all of the circuitry of FIG. 5 may be implemented by one or more virtual machines and/or containers executing on the microprocessor):
acquire the first user image and the second user image (see Azuma Paragraph [0049], The example camera arrays 404 capture a subset of the example user's light field from a plurality of views. The subset of the example user's light field is captured as a plurality of images. The plurality of images represent individual frames from a plurality of videos, where each of the cameras that compose a given camera array 404A records a video. The example camera arrays 404 provide the plurality of images to the example telepresence circuitry 408 via the network 406. The users 402 in the example system 400 utilize respective example camera arrays 404. For example, camera array 404A provides images of user 402A, camera array 404B provides images of user 402B, etc., Paragraph [0175], identify features from a plurality of images, the plurality of images representing a first user and a second user, social view synthesis circuitry to create a first representation of the first user and a second representation of the second user, the representations created using the plurality of images, the representations representing the respective users at specified distances and specified perspectives from a viewer, Paragraph [0182], the plurality of images represent video feeds from a first camera array used by the first user and a second camera array used by the second user, Paragraph [0053], The example telepresence circuitry 408 receives video feeds from the respective users 402 via the network 406. The example telepresence circuitry 408 extracts the subset of the video feed that represents the users 402, Paragraph [0058], The example data receiver circuitry 502 of FIG. 5 receives multiple pluralities of images provided by the camera arrays 404 via the network 406. A given plurality of images represent individual frames of a plurality of videos, where the plurality of videos record a given user from a plurality of views as they communicate. In some examples, the plurality of images are encoded for transmission over the network 406. In some such examples, the example data receiver circuitry 502 decodes the plurality of images and provides the underlying data to the example feature identifier circuitry 504);
execute a 3D display process that displays the first user image and the second user image on the display such that the first user and the second user appear to float off the display surface (see Azuma Paragraph [0183], wherein to create the first representation at specified distances, the social view synthesis circuitry is to further generate a user representation, the user representation to include a three dimensional model or a two dimensional image, create an image of the first user at a first distance from the first camera array, the first distance larger than a physical distance between the first user and the first camera array, and reproject the image of the first user at the first distance onto the user representation, the reprojection to remove distortion caused by the physical distance between the first user and the first camera array, Paragraph [0065], The example shared environment database 506 may also contain one or more data structures to represent the room within a three dimensional (3D) modeling program or 3D rendering software program, Paragraph [0069], The user representation may be a three dimensional model or a two dimensional image. In some examples, the social view synthesis circuitry may determine whether the user representation is a three dimensional or two dimensional image based on the type and computational resources of the displays 410, Paragraph [0092], The data may be used by the augmented headset to recreate the facial expressions and pose of the user representations for a particular frame of input video in three dimensions);
Azuma does not expressively teach
in a fixed mode in which the first user is set as a fixed user, execute the 3D display process such that the first user appears to be closest to the display surface and the second user appears to be in front of the first user.
However, Oz teaches
in a fixed mode in which the first user is set as a fixed user, execute the 3D display process such that the first user appears to be closest to the display surface and the second user appears to be in front of the first user (see Oz Paragraph [0486], Step 2410 may include receiving second participant metadata and first viewpoint metadata by a first unit that may be associated with a first participant, wherein the second participant metadata may be indicative of a pose of a second participant and an expression of the second participant, wherein the first viewpoint metadata may be indicative of a virtual position from which the first participant requests to view an avatar of the second participant, Paragraph [0327], This real-time 3D textured model can then be used to render the view of the user from various angles and camera positions and specifically may be used to correct the viewing position of the virtual camera as if it were virtually located inside the screen of the user—for example—at a virtual location positioned at location that a height and/or lateral location coordinate of at the participant's eyes, Paragraph [0328], The virtual location may be positioned within an imaginary plane that virtually crosses the eyes of the participant—the imaginary plane may (for example) be normal or substantially normal to the display. In this way, a sensation of eye contact may be created for the real time video of the users. The real-time 3D textured model may also be relighted differently than the lighting of the real person in the real environment, to create a more pleasant illumination, e.g., an illumination with less shadowing).
It would have been obvious to one of ordinary skill in the art before the effective filing date of
the claimed invention to combine the teaching of an electronic device that acquires images of two communicating users and displays them as 3D representations on a display (as taught in Azuma), with executing a 3D display process such that users can request another user’s view by selecting a virtual viewing position that simulates direct eye contact (as taught in Oz), the motivation being to reduce the sense of disconnection commonly experienced during virtual communications by restoring spatial presence and the perception of natural eye contact (see Oz Paragraph [0007]).
Regarding Claim 5, it is rejected similarly as Claim 1. The system can be found in Azuma (Abstract, system).
Claims 2 - 4 are rejected under 35 U.S.C. 103 as being unpatentable over Azuma et al. (U.S. Pub. No. 2024/0015263, hereinafter “Azuma”) in view of Oz et al. (U.S. Pub. No. 2023/0171380, hereinafter “Oz”) and Diao et al. (U.S. Pub. No. 2023/0007228, hereinafter “Diao”).
Regarding Claim 2, Azuma in view of Oz teaches
The electronic device according to claim 1, further comprising a camera, wherein
the processing circuitry is further configured to (see Azuma Paragraph [0056], FIG. 5 is a block diagram of an example implementation of example telepresence circuitry to provide remote telepresence communication. The example telepresence circuitry 408 of FIG. 5 may be instantiated (e.g., creating an instance of, bring into being for any length of time, materialize, implement, etc.) by processor circuitry such as a central processing unit executing instructions. Additionally or alternatively, the example telepresence circuitry 408 of FIG. 5 may be instantiated (e.g., creating an instance of, bring into being for any length of time, materialize, implement, etc.) by an ASIC or an FPGA structured to perform operations corresponding to the instructions. It should be understood that some or all of the circuitry of FIG. 5 may, thus, be instantiated at the same or different times. Some or all of the circuitry may be instantiated, for example, in one or more threads executing concurrently on hardware and/or in series on hardware. Moreover, in some examples, some or all of the circuitry of FIG. 5 may be implemented by one or more virtual machines and/or containers executing on the microprocessor):
acquire an image of a third user who is a user of the electronic device, the image being captured by the camera (see Azuma Paragraph [0178], wherein the plurality of images further represents a third user, the social view synthesis circuitry is to further create a third representation of the third user, the third representations created using the plurality of images, the third representations representing the third user at a specified distance and a specified perspectives from a viewer, and the first image further includes the second representation and the third representation at specified locations within the shared environment);
Azuma in view of Oz does not expressively teach
calculate a positional relationship between left and right eyes of the third user and the display surface based on the image of the third user; and
execute the 3D display process based on the positional relationship between the left and right eyes of the third user and the display surface.
However, Diao teaches
calculate a positional relationship between left and right eyes of the third user and the display surface based on the image of the third user (see Diao Paragraph [0071], FIG. 3 schematically shows a geometric relationship model for determining eye space positions. In the embodiment shown in FIG. 3, a geometric relationship model of an XZ plane is shown by taking a direction of the connecting line of the lens centers Oa to Ob of the two cameras as an X-axis direction and taking a direction of the optical axes of the two cameras as a Z-axis direction. In some embodiments, the X-axis direction is also the horizontal direction; the Y-axis direction is also the vertical direction; and the Z-axis direction is a direction perpendicular to the XY plane (also called depth direction), Paragraph [0092], FIG. 5 schematically shows a geometric relationship model for determining eye space positions with a camera and a depth detector. In an embodiment shown in FIG. 5, the camera has a focal length f, an optical axis Z and a focal plane FP; R and L represent the right eye and left eye of the user, respectively; and XR and XL represent X-axis coordinates of imaging of the right eye R and left eye L of the user in the focal plane FP of the camera 155, Paragraph [0091], The shot images are transmitted to the eye positioning image processor 152. The eye positioning image processor may be configured to have a visual identification function (e.g., a face identification function), and may be configured to identify the face based on the shot image and determine eye space positions based on the identified eye positions and the eye depth information of the user, and to determine the viewpoints, where the eyes of the user are located based on the eye space positions. In other embodiments, the 3D processing apparatus determines the viewpoints, where the eyes of the user are located based on the acquired eye space positions); and
execute the 3D display process based on the positional relationship between the left and right eyes of the third user and the display surface (see Diao Abstract, The present disclosure relates to the technical field of 3D display, and discloses a 3D display device, comprising: a multi-viewpoint 3D display screen, which comprises a plurality of composite pixels, wherein each composite pixel of the plurality of composite pixels comprises a plurality of composite subpixels, and each composite subpixel of the plurality of composite subpixels comprises a plurality of subpixels corresponding to a plurality of viewpoints of the 3D display device; a viewing angle determining apparatus, configured to determine a user viewing angle of a user; a 3D processing apparatus, configured to render, based on the user viewing angle, corresponding subpixels of the plurality of composite subpixels according to depth-of-field (DOF) information of a 3D model, Paragraph [0131], Referring to FIG. 9E, two users face the 3D display device; both eyes of a first user are at viewpoints V2 and V4; and both eyes of a second user are at viewpoints V5 and V7. A first 3D image corresponding to a first user viewing angle and a second 3D image corresponding to a second user viewing angle are generated according to the DOF information of the 3D model or the 3D video; left and right eye parallax images corresponding to viewpoints V2 and V4 are generated based on the first 3D image; and left and right eye parallax images corresponding to viewpoints V5 and V7 are generated based on the second 3D image. The 3D processing apparatus renders subpixels, respectively corresponding to the viewpoints V2 and V4 as well as V5 and V7, of the composite subpixels 510, 520 and 530).
It would have been obvious to one of ordinary skill in the art before the effective filing date of
the claimed invention to combine the teaching of an electronic device that acquires images of communicating users, displays them as 3D representations, and executes a 3D display process that enables a user to select a virtual viewing position from which another user appears to be making direct eye contact (as taught in Azuma in view of Oz), with calculating a positional relationship between a user’s eyes and a display surface based on a user’s image to execute a 3D display process (as taught in Diao), the motivation being to create a more accurate and realistic 3D communication by dynamically adjusting and updating the displayed perspective based on user position, thereby improving the viewing experience for users (see Diao Paragraphs [0003], [0004] and [0049]).
Regarding Claim 3, Azuma in view of Oz and Diao teaches
The electronic device according to claim 2, wherein
in the fixed mode, the processing circuitry is further configured to execute the 3D display process such that an apparent position of the first user and and an apparent position of the second user are present on lines connecting between the left and right eyes of the third user and the display surface (see Diao Paragraph [0071], FIG. 3 schematically shows a geometric relationship model for determining eye space positions. In the embodiment shown in FIG. 3, a geometric relationship model of an XZ plane is shown by taking a direction of the connecting line of the lens centers Oa to Ob of the two cameras as an X-axis direction and taking a direction of the optical axes of the two cameras as a Z-axis direction. In some embodiments, the X-axis direction is also the horizontal direction; the Y-axis direction is also the vertical direction; and the Z-axis direction is a direction perpendicular to the XY plane (also called depth direction), Paragraph [0092], FIG. 5 schematically shows a geometric relationship model for determining eye space positions with a camera and a depth detector. In an embodiment shown in FIG. 5, the camera has a focal length f, an optical axis Z and a focal plane FP; R and L represent the right eye and left eye of the user, respectively; and XR and XL represent X-axis coordinates of imaging of the right eye R and left eye L of the user in the focal plane FP of the camera 155, Paragraph [0091], The shot images are transmitted to the eye positioning image processor 152. The eye positioning image processor may be configured to have a visual identification function (e.g., a face identification function), and may be configured to identify the face based on the shot image and determine eye space positions based on the identified eye positions and the eye depth information of the user, and to determine the viewpoints, where the eyes of the user are located based on the eye space positions. In other embodiments, the 3D processing apparatus determines the viewpoints, where the eyes of the user are located based on the acquired eye space positions, Paragraph [0108], the user viewing angle may be an angle of a connecting line between both eyes of the user relative to the coordinate system of the multi-viewpoint 3D display screen or the display plane. In some embodiments, the angle, for example, is an angle θx between the connecting line and the x axis in the coordinate system, or an angle Ay between the connecting line and the y axis in the coordinate system, or is expressed as θ(x,y). In some embodiments, the angle, for example, is an angle between a projection of the connecting line in the xy plane of the coordinate system of the camera and the connecting line. In some embodiments, the angle, for example, is an angle θx between the projection of the connecting line in the xy plane of the coordinate system and the x axis, or an angle θy between the projection of the connecting line in the xy plane of the coordinate system of the camera and the y axis, or is expressed as θ(x,y), Paragraph [0015], In some embodiments, the user viewing angle is an angle between a user sightline and the display plane of the multi-viewpoint 3D display screen, wherein the user sightline is a connecting line between a midpoint of a connecting line, between both eyes of the user, and a center of the multi-viewpoint 3D display screen, Paragraph [0131], Referring to FIG. 9E, two users face the 3D display device; both eyes of a first user are at viewpoints V2 and V4; and both eyes of a second user are at viewpoints V5 and V7. A first 3D image corresponding to a first user viewing angle and a second 3D image corresponding to a second user viewing angle are generated according to the DOF information of the 3D model or the 3D video; left and right eye parallax images corresponding to viewpoints V2 and V4 are generated based on the first 3D image; and left and right eye parallax images corresponding to viewpoints V5 and V7 are generated based on the second 3D image. The 3D processing apparatus renders subpixels, respectively corresponding to the viewpoints V2 and V4 as well as V5 and V7, of the composite subpixels 510, 520 and 530, as applied to viewing positions of users in Azuma in view of Oz).
Regarding Claim 4, Azuma in view of Oz and Diao teaches
The electronic device according to claim 2, wherein
in the fixed mode and when specific information is displayed in a first region of the display surface, the processing circuitry is further configured to execute the 3D display process such that an apparent position of the second user is present on lines connecting between a second region other than the first region of the display surface and the left and right eyes of the third user (see Diao Paragraph [0071], FIG. 3 schematically shows a geometric relationship model for determining eye space positions. In the embodiment shown in FIG. 3, a geometric relationship model of an XZ plane is shown by taking a direction of the connecting line of the lens centers Oa to Ob of the two cameras as an X-axis direction and taking a direction of the optical axes of the two cameras as a Z-axis direction. In some embodiments, the X-axis direction is also the horizontal direction; the Y-axis direction is also the vertical direction; and the Z-axis direction is a direction perpendicular to the XY plane (also called depth direction), Paragraph [0092], FIG. 5 schematically shows a geometric relationship model for determining eye space positions with a camera and a depth detector. In an embodiment shown in FIG. 5, the camera has a focal length f, an optical axis Z and a focal plane FP; R and L represent the right eye and left eye of the user, respectively; and XR and XL represent X-axis coordinates of imaging of the right eye R and left eye L of the user in the focal plane FP of the camera 155, Paragraph [0091], The shot images are transmitted to the eye positioning image processor 152. The eye positioning image processor may be configured to have a visual identification function (e.g., a face identification function), and may be configured to identify the face based on the shot image and determine eye space positions based on the identified eye positions and the eye depth information of the user, and to determine the viewpoints, where the eyes of the user are located based on the eye space positions. In other embodiments, the 3D processing apparatus determines the viewpoints, where the eyes of the user are located based on the acquired eye space positions, Paragraph [0108], the user viewing angle may be an angle of a connecting line between both eyes of the user relative to the coordinate system of the multi-viewpoint 3D display screen or the display plane. In some embodiments, the angle, for example, is an angle θx between the connecting line and the x axis in the coordinate system, or an angle Ay between the connecting line and the y axis in the coordinate system, or is expressed as θ(x,y). In some embodiments, the angle, for example, is an angle between a projection of the connecting line in the xy plane of the coordinate system of the camera and the connecting line. In some embodiments, the angle, for example, is an angle θx between the projection of the connecting line in the xy plane of the coordinate system and the x axis, or an angle θy between the projection of the connecting line in the xy plane of the coordinate system of the camera and the y axis, or is expressed as θ(x,y), Paragraph [0015], In some embodiments, the user viewing angle is an angle between a user sightline and the display plane of the multi-viewpoint 3D display screen, wherein the user sightline is a connecting line between a midpoint of a connecting line, between both eyes of the user, and a center of the multi-viewpoint 3D display screen, Paragraph [0131], Referring to FIG. 9E, two users face the 3D display device; both eyes of a first user are at viewpoints V2 and V4; and both eyes of a second user are at viewpoints V5 and V7. A first 3D image corresponding to a first user viewing angle and a second 3D image corresponding to a second user viewing angle are generated according to the DOF information of the 3D model or the 3D video; left and right eye parallax images corresponding to viewpoints V2 and V4 are generated based on the first 3D image; and left and right eye parallax images corresponding to viewpoints V5 and V7 are generated based on the second 3D image. The 3D processing apparatus renders subpixels, respectively corresponding to the viewpoints V2 and V4 as well as V5 and V7, of the composite subpixels 510, 520 and 530, as applied to viewing positions of users in Azuma in view of Oz).
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure. Refer to PTO-892, Notice of References Cited for a listing of analogous art.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to CARISSA A JONES whose telephone number is (703)756-1677. The examiner can normally be reached Telework M-F 6:30 AM - 4:00 PM CT.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Duc Nguyen can be reached at 5712727503. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/CARISSA A JONES/ Examiner, Art Unit 2691
/DUC NGUYEN/ Supervisory Patent Examiner, Art Unit 2691