Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Priority
Acknowledgement is made of Applicant’s claim of the present applicant claiming priority and benefit under 35 U.S.C. 119(a-d) to Japanese Application No. JP2022-080619 filed 05/17/2022. Acknowledgement is also made of Applicant’s claim of the present application being a 371 of PCT International Application No. PCT/JP2023/011799 filed 03/24/2023.
Information Disclosure Statement
The information disclosure statement (“IDS”) filed on 10/07/2024 was reviewed and the listed references were noted. The information disclosure statement (“IDS”) filed on 06/01/2026 does not include the mandatory IDS Size Fee Assertion; and thus, the references have not been considered.
Drawings
The 7-page drawings have been considered and placed in the file.
Status of Claims
Claims 1-6 are pending.
Claim Rejections - 35 USC § 112
The following is a quotation of 35 U.S.C. 112(b):
(b) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention.
The following is a quotation of 35 U.S.C. 112 (pre-AIA ), second paragraph:
The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the applicant regards as his invention.
Claim 3 is rejected under 35 U.S.C. 112(b) or 35 U.S.C. 112 (pre-AIA ), second paragraph, as being indefinite in that it fails to point out what is included or excluded by the claim language. Claim 3 discloses, “…calculate similarity degrees…”, while claim 2 also discloses similarity degrees. It is unclear whether the similarity degrees of claim 2 and claim 3 are the same similarity degrees. It is suggested that applicant indicate whether the similarity degrees of claim 2 and claim 3 are the same or different by amending. For examination purposes, Examiner interprets the similarity degrees of claim 2 and claim 3 to be different.
Claim Rejections – 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness
rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed
invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows:
1. Determining the scope and contents of the prior art.
2. Ascertaining the differences between the prior art and the claims at issue.
3. Resolving the level of ordinary skill in the pertinent art.
4. Considering objective evidence present in the application indicating obviousness or nonobviousness.
This application currently names joint inventors. In considering patentability of the claims the examiner presumes that the subject matter of the various claims was commonly owned as of the effective filing date of the claimed invention(s) absent any evidence to the contrary. Applicant is advised of the obligation under 37 CFR 1.56 to point out the inventor and effective filing dates of each claim that was not commonly owned as of the effective filing date of the later invention in order for the examiner to consider the applicability of 35 U.S.C. 102(b)(2)(C) for any potential 35 U.S.C. 102(a)(2) prior art against the later invention.
Claims 1-3, and 6 are rejected under 35 U.S.C. 103 as being unpatentable over Yukinori Endo (US 20240071113 A1 w/ EFD of May 20, 2021), in view of Han et al. (“Temporal Alignment Networks for Lon-Term Video”).
Regarding claim 1, Endo teaches, "A video manual generation apparatus comprising: an acquirer configured to acquire: an input-video data set indicative of contents of a task including one or more procedures" (Endo, Para. [0007] discloses; “A video manual generation device in the present disclosure includes processing circuitry to analyze a work procedure manual file in which a work procedure is described and to generate text information data indicating a structure of text included in the work procedure manual file; to analyze a video file of video that is obtained by camera recording of a person executing a work according to the work procedure”) "and one or more input-procedure-text data sets in one-to-one correspondence with the one or more procedures;" (Endo, Para. [0008] discloses, “The video manual generation method includes analyzing a work procedure manual file in which a work procedure is described and generating text information data indicating a structure of text included in the work procedure manual file” Figure 2 of Endo shows the text being in one-to-one correspondence with the procedures.) "an identifier configured to use a task trained model to identify a procedure corresponding to a frame that is any one of a plurality of frames of the input-video data set from among the one or more procedures," (Endo, Paras. [0069] and [0070] disclose; “Subsequently, the link information generation unit 106 generates a link between the “procedure” node and a scene information node (a start time and an end time of the scene) (step S140). The link information generation unit 106 generates links of coordinate positions and movement direction information to the nodes of the three elements of the “procedure” (step S141). FIG. 16 is a diagram showing a generation process of a tree structure of the link information data D106 executed by the link information generation unit 106. FIG. 16 indicates that the link information generation unit 106 forms the tree of the link information data D106 by linking a “procedure” in the tree of the text information data D101 with frame images in the mixed tree of the object information/action information data.”) texts” (Endo, Para. [0008] discloses; “The video manual generation method includes analyzing a work procedure manual file in which a work procedure is described and generating text information data indicating a structure of text included in the work procedure manual file; analyzing a video file of video that is obtained by camera recording of a person executing a work according to the work procedure”) “the video representing the contents of the task constituted of the one or more procedures” (Endo, Para. [0008] discloses; “analyzing a video file of video that is obtained by camera recording of a person executing a work according to the work procedure”) “the one or more texts being in one-to- one correspondence with the one or more procedures” (Endo, Para. [0008] discloses; “The video manual generation method includes analyzing a work procedure manual file in which a work procedure is described and generating text information data indicating a structure of text included in the work procedure manual file” Figure 2 of Endo shows the text being in one-to-one correspondence with the procedures.) “the second information indicating a procedure corresponding to a frame that is any one of a plurality of frames of the video among the one or more procedures;" (Endo, Para. [0070] discloses; “FIG. 16 is a diagram showing a generation process of a tree structure of the link information data D106 executed by the link information generation unit 106. FIG. 16 indicates that the link information generation unit 106 forms the tree of the link information data D106 by linking a “procedure” in the tree of the text information data D101 with frame images in the mixed tree of the object information/action information data”) "and a video manual generator configured to generate video manual data based on the input-video data set and an input-procedure-text data set corresponding to the procedure identified by the identifier from among the one or more input-procedure-text data sets." (Endo, Para. [0008] discloses; “A video manual generation method in the present disclosure is a method executed by a video manual generation device that generates video manual data”). Endo does not explicitly teach, “the task trained model being trained to learn a relationship between first information and second information”. Since Endo does not explicitly disclose this limitation, Examiner relies on the teachings of Han in an analogous field of endeavor. Specifically, Han teaches, “the task trained model being trained to learn a relationship between first information and second information” (Han, Page 3, Section 3 discloses; “In Sec 3.3, we describe a naive training procedure on raw instructional video, with the text video correspondence provided by YouTube ASR, despite the considerable noise.” It would be obvious to combine the trained video manual generation model of Han with the first and second information of Endo to obtain claim 1.)
Endo and Han are both considered to be analogous to the claimed invention because they are in the same field of video manual generation. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have modified Endo to incorporate the teachings of Han in order to train the video generation model to specify a relationship between 2 sets of information. One of ordinary skill in the art would have been motivated to combine the previously described apparatus of Endo with the teachings of Han to ensure the video manual generation model can take the 2 inputs and define a relationship between them. Accordingly, it would have been obvious to combine Endo and Han to obtain the above specified limitations.
Regarding claim 2, the combination of endo and Han teaches, “The video manual generation apparatus according to Claim 1, wherein the task includes a plurality of procedures, wherein the task trained model includes: an image feature model trained to learn a relationship between a frame image and an image feature, the frame image being an image of the frame of the video;” (Han, Page 2, “Visual-Textual Retrieval” discloses; “learns a joint embedding space for both vision and language, either using a dual encoder [2, 19, 24, 25, 31, 34, 48, 53, 54, 55], where visual and textual inputs are independently encoded, or a joint encoder, constructed with multimodal Transformers [13, 40, 44, 45, 66, 67, 81], where vision and text inputs are fed into the cross-modal attention to compute the similarity.” And Han Page 3, “Visual-Textual Backbone” discloses; “Given a long instructional video (e.g. 64s) with its associated sentences, we first extract the visual and textual features with pre-trained networks. Specifically, based on MIL-NCE [47], we use their pre-trained S3D-G backbone to extract video features, and a 2-layer MLP with the word2vec embeddings [49] to extract sentence features.” Examiner interprets the dual encoder to establish a relationship between a frame image and an image feature.) “a natural language feature model trained to learn a relationship between natural languages and natural language features;” (Han, Page 2, “Visual-Textual Retrieval” discloses; “learns a joint embedding space for both vision and language, either using a dual encoder [2, 19, 24, 25, 31, 34, 48, 53, 54, 55], where visual and textual inputs are independently encoded, or a joint encoder, constructed with multimodal Transformers [13, 40, 44, 45, 66, 67, 81], where vision and text inputs are fed into the cross-modal attention to compute the similarity.” And Han Page 3, “Visual-Textual Backbone” discloses; “Given a long instructional video (e.g. 64s) with its associated sentences, we first extract the visual and textual features with pre-trained networks. Specifically, based on MIL-NCE [47], we use their pre-trained S3D-G backbone to extract video features, and a 2-layer MLP with the word2vec embeddings [49] to extract sentence features.” Examiner interprets the dual encoder to establish a relationship between natural language and natural language features.) “a trained model trained to learn a relationship between third information and similarity degrees indicative of a degree of similarity between the frame image and natural languages,” (Han, Page 3, “Problem Scenario” discloses; “A denotes the similarity matrix between frames and the given sentences” It is inherent that there would be a relationship between the similarity and the text and image features.) “the third information being constituted of the image feature and natural language features;” (Han Page 3, “Visual-Textual Backbone” discloses; “Given a long instructional video (e.g. 64s) with its associated sentences, we first extract the visual and textual features with pre-trained networks.) “and a determination model trained to learn a relationship between fourth information and fifth information, the fourth information being constituted of the similarity degrees and the frame corresponding to the frame image” (Han, Page 3, “Problem Scenario” discloses; “A denotes the similarity matrix between frames and the given sentences” and Han, Page 8, “Temporal Alignment on Breakfast-Action” discloses; “Following previous work [9, 10, 18, 41], we report three metrics: frame-wise accuracy (F-Acc), segment wise Intersection-over-Union (IoU) and Intersection-over Detection (IoD).” “Frame-wise accuracy” is interpreted to take the frame number of the frame into account.) “the fifth information being indicative of a procedure corresponding to the similarity degrees and corresponding to the frame corresponding to the frame image among the plurality of procedures” (Han, Page 3, “Problem Scenario” discloses; “A denotes the similarity matrix between frames and the given sentences” and Han Page 4, “Temporal Correspondence” discloses; “The objective is therefore to jointly optimize the visual-textual embedding, such that the similarity score between the sentence and its corresponding visual frames is maximized.”) “wherein the identifier is configured to: use the image feature model to acquire an image feature for the frame that is any one of the plurality of frames of the input-video data set” (Han, Page 3, “Visual-Textual Backbone” discloses; “Given a long instructional video (e.g. 64s) with its associated sentences, we first ex tract the visual and textual features with pre-trained networks.” Since Han is using a video, it is implied that this is performed on a frame of the video.) “use the natural language feature model to acquire natural language features for the input-procedure-text data sets” (Han, Page 3, “Visual-Textual Backbone” discloses; “Given a long instructional video (e.g. 64s) with its associated sentences, we first ex tract the visual and textual features with pre-trained networks.” And Han, Page 3, “Problem Scenario” discloses; “In Sec 3.3, we describe a naive training procedure on raw instructional video, with the text video correspondence provided by YouTube ASR” Examiner interprets the text video correspondence to be input-procedure-text data.) “use the trained model to acquire, for the frame that is any one of the plurality of frames of the input-video data set, similarity degrees corresponding to the acquired image feature and corresponding to the acquired natural language features” (Han, Page 7, “Ablation Study” discloses; “Specifically, we compute the alignment similarity matrix using their textual and visual encoders” Examiner interprets that the encoders produce the image and natural language features and determine a similarity between them.) “and use the determination model to identify, based on the acquired similarity degrees, the procedure corresponding to the frame that is any one of the plurality of frames of the input-video data set from among the plurality of procedures.” (Han, Page 5, “Infer Timestamps” discloses; “To avoid outlier points, for the k-th sentence, we scan its corresponding similarity row by averaging the scores within a temporal window, this window is of the same size as its original YouTube timestamp label, i.e. sentence by the demonstrator. Afterwards, we pick the most confident prediction by taking the argmax. Note that, such operation ends up with a single temporal window with the same duration as the YouTube timestamp. That is to say, we only shift the temporal position of the original YouTube label to its most confident prediction” Examiner interprets this to show that Han identifies the procedure of the frame using a similarity score.) The proposed combination as well as the motivation for combining the Endo and Han references in the rejection of claim 1, apply to claim 2 and are incorporated herein by reference. Thus, the apparatus of claim 2 is met by Endo and Han.
Regarding claim 3, the combination of Endo and Han teaches, “The video manual generation apparatus according to Claim 2, wherein the identifier is configured to: calculate similarity degrees by executing a simple average of, or a weighted average of, similarity degrees obtained by using a current frame of the input-video data set and similarity degrees obtained by using a frame previous to the current frame;” (Han, Page 5, “Infer Timestamps” discloses; “To avoid outlier points, for the k-th sentence, we scan its corresponding similarity row by averaging the scores within a temporal window” Examiner interprets that a “temporal window” could simply be a current and previous frame.) “and use the determination model to acquire, for the frame that is any one of the plurality of frames of the input-video data set, a procedure corresponding to the calculated similarity degrees.” (Han, Page 5, “Infer Timestamps” discloses; “To avoid outlier points, for the k-th sentence, we scan its corresponding similarity row by averaging the scores within a temporal window, this window is of the same size as its original YouTube timestamp label, i.e. sentence by the demonstrator. Afterwards, we pick the most confident prediction by taking the argmax. Note that, such operation ends up with a single temporal window with the same duration as the YouTube timestamp. That is to say, we only shift the temporal position of the original YouTube label to its most confident prediction” Examiner interprets this to show that Han identifies the procedure of the frame using a similarity score.) The proposed combination as well as the motivation for combining the Endo and Han references in the rejection of claim 1, apply to claim 3 and are incorporated herein by reference. Thus, the apparatus of claim 3 is met by Endo and Han.
Regarding claim 6, the combination of Endo and Han teaches, “The video manual generation apparatus according to Claim 1, further comprising a text image generator configured to generate one or more text images in one-to-one correspondence with the one or more procedures based on the one or more input-procedure-text data sets” (Figure 2 of Endo shows a text image (the image in the video display region). This text image is in one-to-one correspondence with a procedure (as shown by the procedure displayed in the workplace procedure manual display region). The procedure is based on the input data set, which can also be seen in the workplace procedure manual display region. (Other support for this mapping can be found in Paras. [0007], [0008], [0048], and [0070]).) “wherein each of the one or more text images represents a corresponding procedure” (Figure 2 of Endo shows that the text image (the image in the video display region) corresponding to a procedure (as shown by the procedure displayed in the workplace procedure manual display region).) “and wherein the video manual generator is configured to generate the video manual data by combining a text image corresponding to the procedure identified by the identifier with the frame image of the input-video data set” (As seen in Figure 2 of Endo, the text image is combined with the frame image of the input video. The whole apparatus of Endo is a video manual generator which is configured to output the display seen in Figure 2.)
Claim 4 is rejected under 35 U.S.C. 103 as being unpatentable over Yukinori Endo, in view of Han et al., in further view of Roychowdhury et al. (US 2023/0154186 A1 w EFD of 11/16/2021).
Regarding claim 4, the combination of Endo and Han teaches, "The video manual generation apparatus according to Claim 2, wherein the determination model is trained to learn, (Han, Page 5, “Infer Timestamps” discloses; “To avoid outlier points, for the k-th sentence, we scan its corresponding similarity row by averaging the scores within a temporal window, this window is of the same size as its original YouTube timestamp label, i.e., sentence by the demonstrator. Afterwards, we pick the most confident prediction by taking the argmax. Note that, such operation ends up with a single temporal window with the same duration as the YouTube timestamp. That is to say, we only shift the temporal position of the original YouTube label to its most confident prediction” Examiner interprets this to show that Han identifies the procedure of the frame using a similarity score.) The combination of Endo and Han does not explicitly teach, "through nonhierarchical clustering". Since the combination of Endo and Han does not explicitly disclose these limitations, Examiner relies on the teachings of Roychowdhury in an analogous field of endeavor. Specifically, Roychowdhury teaches, "through nonhierarchical clustering" (Roychowdhury, Para. [0199] discloses; "the pre-trained embeddings from the supervised action recognition dataset and method are used and a K-means (e.g., K=4) clustering is applied on the embeddings." It would be obvious to use the non-hierarchical clustering technique of K-means clustering disclosed by Roychowdhury to train the apparatus of Endo and Han.).
Endo, Han, and Roychowdhury are considered to be analogous to the claimed invention because they are in the same field of encoding text and video inputs to use for instructional media. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have modified Endo and Han to incorporate the teachings of Roychowdhury in order to train the video generation model using nonhierarchical clustering. One of ordinary skill in the art would have been motivated to combine the previously described apparatus of Endo and Han with the teachings of Roychowdhury to ensure the video manual generation model is trained properly. Accordingly, it would have been obvious to combine Endo, Han, and Roychowdhury to obtain the above specified limitations.
Allowable Subject Matter
Claim 5 is objected to as being dependent upon a rejected base claim, but would be allowable if rewritten in independent for including all of the limitations of the base claim and any intervening claims.
Conclusion
Any inquiry concerning this communication or earlier communications from the examiner should be directed to JUSTIN M. OAKES whose telephone number is (571)272-9379. The examiner can normally be reached 7:30am-5pm.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Amandeep Saini can be reached at (571) 272-3382. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/JUSTIN M OAKES/Examiner, Art Unit 2662
/Siamak Harandi/Primary Examiner, Art Unit 2662