Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Claim Rejections - 35 USC § 101
35 U.S.C. 101 reads as follows:
Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title.
Claim 11 is rejected under 35 U.S.C. 101 because the claimed invention is directed to non-statutory subject matter. The claim does not fall within at least one of the four categories of patent eligible subject matter because the claim 11 recites: “A program for causing a computer to function as a video processing device”. The body of the claim recites computer program steps, such as, “an estimation unit that estimates a 3D skeleton of a subject on the basis of multi-viewpoint images obtained by shooting the subject from a plurality of viewpoints; an application unit that applies a 3D skeleton of the subject estimated by the estimation unit to the subject included in other image different from the multi-viewpoint images and separated from a background in the image; and a generation unit that generates 3D data of a subject to which a 3D skeleton is applied by the application unit.”, which are nothing more than just programmed instructions to be performed by the system.
Therefore, the steps/elements recited in claim 11 are non-statutory. Similarly, computer programs claimed as computer listings per se, i.e., the descriptions or expressions of the programs, are not physical “things”. They are neither computer components nor statutory processes, as they are not “acts” being performed. Such claimed computer programs do not define any structural and functional interrelationships between the computer program and other claimed elements of a computer which permit the computer program’s functionality to be realized. In contrast, a claimed non-transitory computer-readable medium encoded with a computer program is a computer element which defines structural and functional interrelationships between the computer program and the rest of the computer which permit the computer program’s functionality to be realized, and is thus statutory. Accordingly, it is important to distinguish claims that define descriptive material per se from claims that define statutory inventions.
Claim Interpretation
The following is a quotation of 35 U.S.C. 112(f):
(f) Element in Claim for a Combination. – An element in a claim for a combination may be expressed as a means or step for performing a specified function without the recital of structure, material, or acts in support thereof, and such claim shall be construed to cover the corresponding structure, material, or acts described in the specification and equivalents thereof.
The following is a quotation of pre-AIA 35 U.S.C. 112, sixth paragraph:
An element in a claim for a combination may be expressed as a means or step for performing a specified function without the recital of structure, material, or acts in support thereof, and such claim shall be construed to cover the corresponding structure, material, or acts described in the specification and equivalents thereof.
The claims in this application are given their broadest reasonable interpretation using the plain meaning of the claim language in light of the specification as it would be understood by one of ordinary skill in the art. The broadest reasonable interpretation of a claim element (also commonly referred to as a claim limitation) is limited by the description in the specification when 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is invoked.
As explained in MPEP § 2181, subsection I, claim limitations that meet the following three-prong test will be interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph:
(A) the claim limitation uses the term “means” or “step” or a term used as a substitute for “means” that is a generic placeholder (also called a nonce term or a non-structural term having no specific structural meaning) for performing the claimed function;
(B) the term “means” or “step” or the generic placeholder is modified by functional language, typically, but not always linked by the transition word “for” (e.g., “means for”) or another linking word or phrase, such as “configured to” or “so that”; and
(C) the term “means” or “step” or the generic placeholder is not modified by sufficient structure, material, or acts for performing the claimed function.
Use of the word “means” (or “step”) in a claim with functional language creates a rebuttable presumption that the claim limitation is to be treated in accordance with 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph. The presumption that the claim limitation is interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is rebutted when the claim limitation recites sufficient structure, material, or acts to entirely perform the recited function.
Absence of the word “means” (or “step”) in a claim creates a rebuttable presumption that the claim limitation is not to be treated in accordance with 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph. The presumption that the claim limitation is not interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is rebutted when the claim limitation recites function without reciting sufficient structure, material or acts to entirely perform the recited function.
Claim limitations in this application that use the word “means” (or “step”) are being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, except as otherwise indicated in an Office action. Conversely, claim limitations in this application that do not use the word “means” (or “step”) are not being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, except as otherwise indicated in an Office action.
This application includes one or more claim limitations that do not use the word “means,” but are nonetheless being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, because the claim limitation(s) uses a generic placeholder that is coupled with functional language without reciting sufficient structure to perform the recited function and the generic placeholder is not preceded by a structural modifier. Such claim limitations are: “an estimation unit”, “an application unit”, and “a generation unit” in claims 1 and 11.
Because these claim limitations is/are being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, they are being interpreted to cover the corresponding structure described in the specification as performing the claimed function, and equivalents thereof.
If applicant does not intend to have these limitations interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, applicant may: (1) amend the claim limitation(s) to avoid it/them being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph (e.g., by reciting sufficient structure to perform the claimed function); or (2) present a sufficient showing that the claim limitations recite sufficient structure to perform the claimed function so as to avoid it/them being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph.
Claim Rejections - 35 USC § 112
The following is a quotation of 35 U.S.C. 112(b):
(b) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention.
The following is a quotation of 35 U.S.C. 112 (pre-AIA ), second paragraph:
The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the applicant regards as his invention.
Claims 2-5 are rejected under 35 U.S.C. 112(b) or 35 U.S.C. 112 (pre-AIA ), second paragraph, as being indefinite for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor (or for applications subject to pre-AIA 35 U.S.C. 112, the applicant), regards as the invention.
Regarding claim 2, claim 2 recites “a 3D skeleton of the subject” in lines 4-5. As claim 1 also recites “a 3D skeleton of the subject” in lines 5-6, it is indeterminate whether “a 3D skeleton of the subject” recited in claim 2 is the same “a 3D skeleton of the subject” in claim 1. If this is the case, then “a 3D skeleton of the subject” in claim 2 should be “the 3D skeleton of the subject”. However, if this is not the case, “a 3D skeleton of the subject” in claims 1 and 2 should be “a first 3D skeleton of the subject” and “a second 3D skeleton of the subject” respectively. To advance compact prosecution, the Examiner interprets “a 3D skeleton of the subject” in claim 2 as “the 3D skeleton of the subject”.
Regarding claim 3, claim 3 recites “a 3D skeleton of the subject in a plurality of frames” in lines 3-4 and “a 3D skeleton of the subject in the plurality of frames” in lines 7-8. As “the plurality of frames” seems to relates to “a plurality of frames”, it is indeterminate if “a 3D skeleton of the subject in the plurality of frames” in lines 7-8 is the same “a 3D skeleton of the subject in a plurality of frames” in lines 3-4 or not. If they are the same, “a 3D skeleton of the subject in the plurality of frames” should be “the 3D skeleton of the subject in the plurality of frames”. If this is not the case, these limitations should be “a first 3D skeleton of the subject in a plurality of frames” for lines 3-4 and “a second 3D skeleton of the subject in the plurality of frames” for lines 7-8. To advance compact prosecution, the Examiner interprets “a 3D skeleton of the subject in the plurality of frames” in lines 7-8 as the same 3D skeleton as “a 3D skeleton of the subject in a plurality of frames”.
Similar to claim 2, claim 4 recites the same “a 3D skeleton of the subject”, therefore, also rejected under 112(b) for indefiniteness. The same suggestions are provided for claim 4 as well. To advance compact prosecution, the Examiner interprets “a 3D skeleton of the subject” in claim 4 as “the 3D skeleton of the subject”.
Dependent claim 5 incorporate independent claim deficiency from claim 4 and thus also rejected under 112(b).
Claim Rejections - 35 USC § 102
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action:
A person shall be entitled to a patent unless –
(a)(1) the claimed invention was patented, described in a printed publication, or in public use, on sale, or otherwise available to the public before the effective filing date of the claimed invention.
Claims 1-3, 6, and 10-11 are rejected under 35 U.S.C. 102(a)(1) as being anticipated by Masatoshi (JP 2020065229 A).
Regarding claim 1, Masatoshi teaches a video processing device comprising: ([0011] “The video communication device according to the present invention is characterized by comprising”)
an estimation unit that estimates a 3D skeleton of a subject on the basis of multi-viewpoint images obtained by shooting the subject from a plurality of viewpoints; ([0075] “The receiving unit 134 receives the user's skeletal position determined from the video by the image recognition server 60.”, where “The image recognition server 60 estimates the skeletal position of each user at each location based on the video footage received from each location.” [0074] and “The transmission unit 133 transmits video footage captured by multiple cameras 22A and 22B at each location to the image recognition server 60.” [0073])
an application unit that applies a 3D skeleton of the subject estimated by the estimation unit to the subject included in other image different from the multi-viewpoint images and separated from a background in the image; ([0011] “a model control unit that changes the posture of a rigged 3D model having the appearance of the communication partner based on the skeletal position information”, where [0077] “The control unit 131 reads the 3D real avatars of users at other locations from the storage unit 135 and reflects the received skeletal position onto the 3D real avatars.”)
and a generation unit that generates 3D data of a subject to which a 3D skeleton is applied by the application unit ([0011] “and a display unit that displays the 3D model superimposed on a real landscape”, where “The display unit 132 places a 3D real avatar reflecting the skeletal position for each user at other locations at the user's coordinates in the spatially shared coordinate system, and uses the coordinates in the spatially shared coordinate system of the user wearing the AR devices 23A and 23B as the viewpoint, and renders the 3D real avatar in the line of sight direction of the AR devices 23A and 23B.” [0079]).
Regarding claim 2, Masatoshi teaches the video processing device according to claim 1, wherein on the basis of a rig attached to the subject included in the other image, the application unit applies a 3D skeleton of the subject to the subject ([0020] “3D Real Avatars are equipped with a rig that moves the 3D model in accordance with the movement of the human skeleton. Therefore, if the user's skeletal position is known, it is possible to make the avatar perform the same movements as the user by manipulating the rig based on that skeletal position.”).
Regarding claim 3, Masatoshi teaches the video processing device according to claim 1, wherein the estimation unit estimates a 3D skeleton of the subject in a plurality of frames on the basis of a plurality of the multi-viewpoint images continuously captured ([0060] “At each of the locations 1, 2, and 3, the coordinates of each location for users A-F are determined from the video footage captured by cameras 22A and 22B (step S21), and the coordinates of each location for users A-F are transmitted to the coordinate management server 50 (step S22)”, where “From this point onward, each of the locations 1, 2, and 3 continues to calculate and transmit the coordinates of users A-F at each location from the video footage, while simultaneously receiving the coordinates of users A-F at the other locations 1, 2, and 3 and converting them into coordinates in the shared spatial coordinate system.” [0071]).
Regarding claim 6, Masatoshi teaches the video processing device according to claim 1, wherein on the basis of the multi-viewpoint images including a plurality of subjects, the estimation unit estimates a 3D skeleton of each of the plurality of subjects, (Figure 4, where “The 3D real avatar management server 40 manages the 3D real avatars of users A-F.” [0034], “The image recognition server 60 receives video footage of users A-F from each location and estimates the skeletal positions of users A-F.” [0036], and “The coordinates and skeletal positions of users A-F are reflected in the 3D realistic avatars.” [0031])
the application unit applies a 3D skeleton of each of the plurality of subjects to each of the subjects included in other images different from the multi-viewpoint images, and the generation unit generates 3D data of each subject to which a 3D skeleton is applied by the application unit. ([0040] “The avatar processing unit 13 receives the skeletal position of users at other locations from the image recognition server 60, reflects the skeletal position in the 3D real avatar, and causes the 3D real avatar to represent the actions of users at other locations.”)
Regarding claim 10, claim 10 recites substantially similar limitations to claim 1, but in a method form. Masatoshi further teaches a video processing method including execution by a computer to: ([0010] “The video communication method according to the present invention is characterized by comprising the steps of: receiving skeletal position information of a communication partner detected based on video footage taken from multiple directions; changing the posture of a rigged 3D model having the appearance of the communication partner based on the skeletal position information; and displaying the 3D model superimposed on a real landscape.”)
Regarding claim 11, claim 11 recites substantially similar limitations to claim 1, but in a program form. Masatoshi further teaches a program for causing a computer to function as a video processing device ([0012] “The video communication program according to the present invention is characterized by causing a computer to execute the above-described video communication method”).
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claim 4 is rejected under 35 U.S.C. 103 as being unpatentable over Masatoshi, in view of Asayama (U.S. 2023/0316543 A1).
Masatoshi teaches the video processing device according to claim 1, but fails to teach wherein in a case where second 3D data generated in advance has a defect, the estimation unit estimates a 3D skeleton of a subject in a frame of a multi-viewpoint image corresponding to the second 3D data, the application unit applies a 3D skeleton of the subject estimated by the estimation unit to a subject in a frame of a multi-viewpoint image corresponding to the second 3D data, and the generation unit generates 3D data of a subject to which a 3D skeleton is applied by the application unit instead of the second 3D data. However, this is known in the art as taught by Asayama.
Asayama teaches wherein in a case where second 3D data generated in advance has a defect, [0066] “The skeleton estimation device 100 detects the presence or absence of the defective skeleton with respect to the estimated skeleton information (CG data K1 to K11).”
the estimation unit estimates a 3D skeleton of a subject in a frame of a multi-viewpoint image corresponding to the second 3D data, the application unit applies a 3D skeleton of the subject estimated by the estimation unit to a subject in a frame of a multi-viewpoint image corresponding to the second 3D data, and the generation unit generates 3D data of a subject to which a 3D skeleton is applied by the application unit instead of the second 3D data ([0069] “The skeleton estimation device 100 becomes able to modify the skeleton information of the defective skeleton NG part by performing the skeleton estimation again for the skeleton information (CG data K3 to K6) of the detected defective skeleton NG.”, where “The skeleton estimation device 100 performs the training-type skeleton recognition for the depth image including the distance information of each of the X, Y, and Z axes as the three-dimensional skeleton information of the gymnast acquired from the 3D laser sensor, calculates the 3D joint coordinates, and fits the 3D joint coordinates to a human body model. The CG data K1 to K 11 in FIG. 5B are fitted estimated skeleton information.” [0066]).
Asayama is analogous to the claimed invention, as both relate to 3D skeleton estimation from multi-view images. Asayama further teaches that “Furthermore, it becomes possible to perform the skeleton determination again for the portion of the detected defective skeleton, and it becomes possible to improve accuracy of the estimated skeleton over the entire predetermined period.” [0064]. Therefore, it would be obvious for one of ordinary skill in the art before the effective filing date of the claimed invention to incorporate the teachings of Asayama to Masatoshi to improve accuracy of the estimated skeleton.
Claim 5 is rejected under 35 U.S.C. 103 as being unpatentable over Masatoshi, in view of Asayama, and further in view of Seyrek Pierre et al. (U.S. 2024/0029347 A1, hereinafter Seyrek).
Regarding claim 5, the combination of Masatoshi and Asayama teaches the video processing device according to claim 4, but fails to teach wherein the generation unit generates third 3D data by taking a logical sum of the second 3D data and 3D data of the subject. However, this is known in the art as taught by Seyrek.
Seyrek teaches that generat[ing] third 3D data by taking a logical sum of the second 3D data and 3D data of the subject, as “By combining two or more images it is possible to derive 3D coordinates for each key point that is visible in at least two images. If the animal is too long to be captured in its entirety by the camera setup, it may be possible to generate the entire 3D skeleton based on composite images generated by images with overlapping regions.” [0065] (Note: the Specifications describe logical sum as “logical sum (OR operation) of both the frames and determine, as the subject, a part including the data of the subject in any of the frames.” [0079]). Seyrek is also analogous to the claimed invention, as both relate to creating 3D skeletons from multiview images. Therefore, it would be obvious for one of ordinary skill in the art before the effective filing date of the claimed invention to incorporate the teachings of Seyrek to the combination of Masatoshi and Asayama in order to generate 3D data in its entirety from multiple images.
Claims 7 and 8 are rejected under 35 U.S.C. 103 as being unpatentable over Masatoshi, in view of Kim et al. (U.S. Patent No. 8,766,977, hereinafter Kim).
Regarding claim 7, Masatoshi teaches the video processing device according to claim 6, but fails to teach wherein, in a case of applying a 3D skeleton of each of the plurality of subjects to each subject, the application unit specifies data of an application destination subject by taking a logical product of data corresponding to the subject separated from a background between a plurality of images including at least one subject. However, this is known in the art as taught by Kim.
Kim teaches wherein, in a case of applying a 3D skeleton of each of the plurality of subjects to each subject, the application unit specifies data of an application destination subject by taking a logical product of data corresponding to the subject separated from a background between a plurality of images including at least one subject ([col. 4, lines 19-28] “the skeleton model generation unit 104 may generate the 3D skeleton model of the entire body of the user using an error range of each of the 3D skeleton models extracted from the silhouette and from the position or orientation of the joint of the user. In this instance, the skeleton model generation unit 104 may combine the error range of each of the 3D skeleton models extracted from the silhouette and the position or orientation of the joint of the user, and select a position having a minimum error.”, where “Referring to operation (IV) of FIG. 2, and FIG. 1, the apparatus 100 may extract a silhouette of the user from the image data. When the image sensor 107 obtains the multi-view image data again, the silhouette may be extracted per point of view.” [col. 4, lines 58-62]. Note: the Specifications describe logical product as “logical product (AND processing) of parts determined as silhouettes in the frame 270 and the frame 272. Then, since only the parts determined to be silhouettes in both frames are extracted, the video processing device 100 can specify an appropriate silhouette of one person like a frame 276.” [0112]).
Kim is analogous to the claimed invention, as both relate to generating 3D skeletons from multiview images. Kim further teaches that “According to example embodiments, the error range of each of the 3D skeleton models extracted from the silhouette and from the position or orientation of the joint of the user may be used, thereby generating the 3D skeleton model having a high reliability and high accuracy.” [col. 10, lines 29-33]. Therefore, it would be obvious for one of ordinary skill in the art before the effective filing date of the claimed invention to incorporate the teachings of Kim to Masatoshi in order to generate highly reliable and accurate skeleton models.
Regarding claim 8, the combination of Masatoshi and Kim teaches the video processing device according to claim 7 and a 3D skeleton after correction, but does not explicitly teach wherein the generation unit corrects a shape of the subject specified to generate 3D data of the subject on the basis of a 3D skeleton after correction. However, the combination of Masatoshi and Kim teaches that it is known in the art for the shape of the 3D data to be generated on the basis of a 3D skeleton (Masatoshi; [0020] “The 3D realistic avatar generation server 30 deforms a humanoid base model that has been rigged for its entire body to match the shape of the 3D scan data and generates a 3D realistic avatar.”), which would not be limited to the corrected 3D skeleton as taught by Kim.
Claim 9 is rejected under 35 U.S.C. 103 as being unpatentable over Masatoshi, in view of Seo et al. (U.S. Patent No. 8,766,977, hereinafter Seo).
Masatoshi teaches the video processing device according to claim 6, but fails to teach wherein the generation unit generates 3D data in which the respective subjects to which a 3D skeleton is applied by the application unit are integrated into one. However, this is known in the art as taught by Seo.
Seo teaches wherein the generation unit generates 3D data in which the respective subjects to which a 3D skeleton is applied by the application unit are integrated into one ([0053] “Here, the goal is to generate a 3D skeleton, so we obtain a partial skeleton of the object captured at each location (DL-based 3D Skeleton Generation), align these partial skeletons (Skeleton Alignment), and integrate them (Skeleton Integration) to obtain a complete 3D skeleton.”)
Seo is analogous to the claimed invention, as both relate to generating a 3D skeleton using multi-viewpoint images. Seo further teaches that “by calculating parameters using the joints of a partial skeleton of a multi-view RGB-D image as feature points and integrating the skeleton based thereon, a 3D skeleton with high reliability can be generated, and the generated 3D skeleton can accurately express the shape and movement of 3D volume information without temporal interruption.” [0024]. Therefore, it would be obvious for one of ordinary skill in the art before the effective filing date of the claimed invention before the effective filing date of the claimed invention to incorporate the teachings of Seo to Masatoshi in order to reliably generate an accurate 3D skeleton without interruptions.
Conclusion
Any inquiry concerning this communication or earlier communications from the examiner should be directed to ALICIA HA whose telephone number is (571)272-3601. The examiner can normally be reached Mon-Thurs 9:30 AM - 6:30 PM, and Fri 9:30 AM - 1:30 PM.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Kee Tung can be reached at (571) 272-7794. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/ALICIA HA/Examiner, Art Unit 2611
/KEE M TUNG/Supervisory Patent Examiner, Art Unit 2611