DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Claim Objections
Claims 1 and 6 are objected to because of the following informalities: the language “to capturing of” is not correct. Appropriate correction is required.
Claim Rejections - 35 USC § 102
The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action:
A person shall be entitled to a patent unless –
(a)(2) the claimed invention was described in a patent issued under section 151, or in an application for patent published or deemed published under section 122(b), in which the patent or application, as the case may be, names another inventor and was effectively filed before the effective filing date of the claimed invention.
Claims 1-2 and 6-8 are rejected under 35 U.S.C. 102(a)(2) as being anticipated by Takehara (US 20240221240 A1).
Regarding to claim 1, Takehara discloses a method implemented by a server (Fig. 2; [0017]: a display system 1 includes the display device 10 and an information processing device 12; [0018]: the information processing device 12 is a device that performs information processing on information for the virtual space SV; the information processing device 12 is referred to as a server that transmits and receives information to and from the display device 10 of the user U and performs image processing on images for the virtual space SV), the method comprising:
receiving data related to capturing of an object by a capturing device (Fig. 2; [0018]: a server transmits and receives information to and from the display device 10 of the user U and performs image processing on images for the virtual space SV; [0022]: detect the surroundings of the display device 10 (user U) using a 3D camera, such as a stereo camera; camera captures the image of the real space SR);
obtaining position information indicating a first position of the object from the data received ([0022]: the real space detection unit 28 is a sensor that detects a position of the user U, a posture of the user U, and the surroundings of the display device 10, i.e. user U, in the real space SR; [0024]: the real object information acquisition unit 40 acquires information on a real object OR; Fig. 7; [0045]: the real object information acquisition unit 40 acquires information on a real object OR via the real space detection unit 28); and
outputting information for displaying an image corresponding to the object at a second position in a virtual space (Fig. 1; [0016]: the display device 10 displays an image for the virtual space SV according to a motion of the user U in the real space SR; [0029]: the virtual object information acquisition unit 42 acquires information on a virtual object OV; the virtual object OV is a virtual object that exists in the virtual space SV and is a subject that is displayed as an image for the virtual space SV in the coordinate system in the virtual space SV), the second position corresponding to the first position indicated by the position information obtained (Fig. 1; [0016]: the display device 10 displays an image for the virtual space SV according to a motion of the user U in the real space SR; [0029]: the virtual object information acquisition unit 42 acquires the position, i.e. coordinates, of the virtual object OV with respect to the coordinates of the origin of the virtual space SV as the positional information on the virtual object OV; [0034]: the user position acquisition unit 44 acquires user positional information; the user positional information is information indicating the position of the user UV in the virtual space SV that is set based on the information on the real object OR that is acquired by the real object information acquisition unit 40 and the information on the virtual object OV that is acquired by the virtual object information acquisition unit 42).
Regarding to claim 2, Takehara discloses the method according to claim 1, wherein
the obtaining includes obtaining the position information by performing image processing on an image obtained by the capturing device capturing the object (Takehara; [0022]: the real space detection unit 28 is a sensor that detects a position of the user U, a posture of the user U, and the surroundings of the display device 10, i.e. user U, in the real space SR; [0024]: the real object information acquisition unit 40 acquires information on a real object OR; Fig. 7; [0045]: the real object information acquisition unit 40 acquires information on a real object OR via the real space detection unit 28), and
the outputting includes calculating the second position using a position of the capturing device and the position information obtained (Takehara; Fig. 1; [0016]: the display device 10 displays an image for the virtual space SV according to a motion of the user U in the real space SR; [0031]: the virtual object information acquisition unit 42 extracts the virtual object OV based on the shape information on the real object OR and the shape information on the virtual object OV; Fig. 7; [0045]: the virtual object information acquisition unit 42 acquires information on a virtual object OV from the information processing device 12), and outputting the information for displaying the image at the second position calculated (Fig. 1; [0016]: the display device 10 displays an image for the virtual space SV according to a motion of the user U in the real space SR; Fig. 7; [0045]: the user position acquisition unit 44 sets user positional information based on the information on the real object OR and the information on the virtual object OV (step S16); the providing unit 46 sets a position that the user positional information presents for the position of a user UV in a virtual space SV and provides the virtual space SV to the user U).
Regarding to claim 6, Takehara discloses a server (Fig. 2; [0017]: a display system 1 includes the display device 10 and an information processing device 12; [0018]: the information processing device 12 is a device that performs information processing on information for the virtual space SV; the information processing device 12 is referred to as a server that transmits and receives information to and from the display device 10 of the user U and performs image processing on images for the virtual space SV) comprising:
a communicator (Fig. 3; [0021]: the display device 10 communicates with an external device, such as the information processing device 12, by wireless communication); and
a processor ([0018]: a computer includes a computing device including computing circuitry, such as a CPU, i.e. Central Processing Unit, and a storage unit and executes processing by reading and executing a program, i.e., software, from the storage.), wherein
the communicator receives data related to capturing of an object by a capturing device ( Fig. 2; [0018]: a server transmits and receives information to and from the display device 10 of the user U and performs image processing on images for the virtual space SV; [0022]: detect the surroundings of the display device 10 (user U) using a 3D camera, such as a stereo camera; camera captures the image of the real space SR), and the processor obtains position information indicating a first position of the object from the data received by the communicator ([0022]: the real space detection unit 28 is a sensor that detects a position of the user U, a posture of the user U, and the surroundings of the display device 10, i.e. user U, in the real space SR; [0024]: the real object information acquisition unit 40 acquires information on a real object OR), and outputs information for displaying an image corresponding to the object at a second position in a virtual space (Fig. 1; [0016]: the display device 10 displays an image for the virtual space SV according to a motion of the user U in the real space SR; [0029]: the virtual object information acquisition unit 42 acquires information on a virtual object OV; the virtual object OV is a virtual object that exists in the virtual space SV and is a subject that is displayed as an image for the virtual space SV in the coordinate system in the virtual space SV), the second position corresponding to the first position indicated by the position information obtained (Fig. 1; [0016]: the display device 10 displays an image for the virtual space SV according to a motion of the user U in the real space SR; [0029]: the virtual object information acquisition unit 42 acquires the position, i.e. coordinates, of the virtual object OV with respect to the coordinates of the origin of the virtual space SV as the positional information on the virtual object OV; [0034]: the user position acquisition unit 44 acquires user positional information; the user positional information is information indicating the position of the user UV in the virtual space SV that is set based on the information on the real object OR that is acquired by the real object information acquisition unit 40 and the information on the virtual object OV that is acquired by the virtual object information acquisition unit 42).
Regarding to claim 7, Takehara discloses a method implemented by a capturing device (Fig. 2; [0017]: a display system 1 includes the display device 10 and an information processing device 12; [0018]: the information processing device 12 is a device that performs information processing on information for the virtual space SV; the information processing device 12 is referred to as a server that transmits and receives information to and from the display device 10 of the user U and performs image processing on images for the virtual space SV), the method comprising:
capturing an object ([0022]: detect the surroundings of the display device 10 (user U) using a 3D camera, such as a stereo camera; camera captures the image of the real space SR); and
by transmitting data related to the capturing of the object to a server, causing the server to (Fig. 2; [0018]: the information processing device 12 is a device that performs information processing on information for the virtual space SV; the information processing device 12 is referred to as a server that transmits and receives information to and from the display device 10 of the user U and performs image processing on images for the virtual space SV; Fig. 3; [0021]: the display device 10 communicates with an external device, such as the information processing device 12, by wireless communication; [0022]: detect the surroundings of the display device 10 (user U) using a 3D camera, such as a stereo camera; camera captures the image of the real space SR): obtain position information indicating a first position of the object captured by the capturing device ([0022]: the real space detection unit 28 is a sensor that detects a position of the user U, a posture of the user U, and the surroundings of the display device 10, i.e. user U, in the real space SR; [0024]: the real object information acquisition unit 40 acquires information on a real object OR); and output information for displaying an image corresponding to the object at a second position in a virtual space (Fig. 1; [0016]: the display device 10 displays an image for the virtual space SV according to a motion of the user U in the real space SR; [0029]: the virtual object information acquisition unit 42 acquires information on a virtual object OV; the virtual object OV is a virtual object that exists in the virtual space SV and is a subject that is displayed as an image for the virtual space SV in the coordinate system in the virtual space SV), the second position corresponding to the first position indicated by the position information obtained (Fig. 1; [0016]: the display device 10 displays an image for the virtual space SV according to a motion of the user U in the real space SR; [0029]: the virtual object information acquisition unit 42 acquires the position, i.e. coordinates, of the virtual object OV with respect to the coordinates of the origin of the virtual space SV as the positional information on the virtual object OV; [0034]: the user position acquisition unit 44 acquires user positional information; the user positional information is information indicating the position of the user UV in the virtual space SV that is set based on the information on the real object OR that is acquired by the real object information acquisition unit 40 and the information on the virtual object OV that is acquired by the virtual object information acquisition unit 42).
Regarding to claim 8, Takehara discloses a capturing device (Fig. 2; [0017]: a display system 1 includes the display device 10 and an information processing device 12; [0018]: the information processing device 12 is a device that performs information processing on information for the virtual space SV; the information processing device 12 is referred to as a server that transmits and receives information to and from the display device 10 of the user U and performs image processing on images for the virtual space SV) comprising:
a communicator (Fig. 3; [0021]: the display device 10 communicates with an external device, such as the information processing device 12, by wireless communication);
a processor ([0018]: a computer includes a computing device including computing circuitry, such as a CPU, i.e. Central Processing Unit, and a storage unit and executes processing by reading and executing a program, i.e., software, from the storage); and
a capturer ([0022]: detect the surroundings of the display device 10 (user U) using a 3D camera, such as a stereo camera; camera captures the image of the real space SR), wherein
the capturer captures an object ([0022]: detect the surroundings of the display device 10 (user U) using a 3D camera, such as a stereo camera; camera captures the image of the real space SR), and
the rest claim limitations are similar to claim limitations recited in claim 7. Therefore, same rational used to reject claim 7 is also used to reject rest claim limitations.
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claims 3-5 are rejected under 35 U.S.C. 103 as being unpatentable over Takehara (US 20240221240 A1) and in view of Kar (US 20200228774 A1).
Regarding to claim 3, Takehara discloses the method according to claim 1, further comprising:
obtaining an image obtained by the capturing device capturing the object (Takehara; [0022]: detect the surroundings of the display device 10, i.e. user U, using a 3D camera, such as a stereo camera; camera captures the image of the real space SR), wherein
Takehara fails to explicitly disclose:
the outputting includes outputting, as the information, an image output by a trained model in response to inputting at least the image obtained by capturing the object into the trained model, and
the trained model is generated by machine learning to generate and output a new image using an input image.
In same field of endeavor, Kar teaches:
the outputting includes outputting, as the information, an image output by a trained model in response to inputting at least the image obtained by capturing the object into the trained model ([0060]: a deep learning pipeline is trained on renderings of natural scenes and the use of an intermediate volumetric scene representation; [0113]: the captured image is promoted to a local multiplane image by a trained CNN; [0115]: the captured image is promoted to an MPI by applying a convolutional neural network (CNN) to the focal image; Fig. 34; [0289]: train a novel view model; [0147]: this baseline involves training a network uses the same 3D CNN architecture as the MPI prediction network; uses a second 2D CNN to composite these warped input images into a single output rendered view; [0290]: training is performed by generating novel views of 3D models; those rendered images are then used to generate a novel view from a target viewpoint; [0291]: a view synthesis model is trained differently for different contexts; [0298]: the model is trained on the final blended rendering because the fixed rendering and blending functions are differentiable; [0302]: the model may continue to be trained until one or more stopping criteria are met. For example, training may continue until the marginal change in the model between successive iterations falls below a threshold; Fig. 34; [0303]: the trained model is stored at 3422), and
the trained model is generated by machine learning to generate and output a new image using an input image ([0060]: a deep learning pipeline is trained on renderings of natural scenes and the use of an intermediate volumetric scene representation; [0075]: the high quality of the deep learning predicted local scene representations allows the synthesis of superior renderings without requiring aggregating geometry estimates over large view neighborhoods; [0114]: the input camera poses may be estimated and an MPI predicted for each input view using a trained neural network; the deep learning pipeline may be used to predict an MPI for each input sampled view; [0147]: this baseline involves training a network uses the same 3D CNN architecture as the MPI prediction network; uses a second 2D CNN to composite these warped input images into a single output rendered view; Fig. 6; [0189]: one or more multi-view interactive digital media representations is generated from the content and context models at 608).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify Takehara to include the outputting includes outputting, as the information, an image output by a trained model in response to inputting at least the image obtained by capturing the object into the trained model, and the trained model is generated by machine learning to generate and output a new image using an input image as taught by Kar. The motivation for doing so would have been to employ a sampling framework in combination with deep-learning-based view synthesis to significantly decrease the dense sampling requirements of traditional light field rendering; to produce improved renderings; to present a user with an interactive and immersive active viewing experience; to train a novel view model, as taught by Kar in paragraphs [0070], [0077], [0161], and [0289-291],
Regarding to claim 4, Takehara in view of Kar discloses the method according to claim 3, wherein
the outputting includes outputting, as the information, image data output by the trained model in response to inputting into a trained model the image obtained by capturing the object and textual information related to an image output by the trained model (Kar; [0060]: a deep learning pipeline is trained on renderings of natural scenes and the use of an intermediate volumetric scene representation; [0075]: the high quality of the deep learning predicted local scene representations allows the synthesis of superior renderings without requiring aggregating geometry estimates over large view neighborhoods; [0114]: the input camera poses may be estimated and an MPI predicted for each input view using a trained neural network; the deep learning pipeline may be used to predict an MPI for each input sampled view).
Same motivation of claim 3 is applied here
Regarding to claim 5, Takehara in view of Kar discloses the method according to claim 3, wherein
the trained model is generated by the machine learning using a plurality of datasets possessed by a plurality of operators (Kar; [0060]: a deep learning pipeline is trained on renderings of natural scenes and the use of an intermediate volumetric scene representation; [0075]: the high quality of the deep learning predicted local scene representations allows the synthesis of superior renderings without requiring aggregating geometry estimates over large view neighborhoods; [0114]: the input camera poses may be estimated and an MPI predicted for each input view using a trained neural network; the deep learning pipeline may be used to predict an MPI for each input sampled view).
Same motivation of claim 3 is applied here.
Conclusion
Any inquiry concerning this communication or earlier communications from the examiner should be directed to Hai Tao Sun whose telephone number is (571)272-5630. The examiner can normally be reached 9:00AM-6:00PM.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Daniel Hajnik can be reached at 5712727642. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/HAI TAO SUN/Primary Examiner, Art Unit 2616