DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Information Disclosure Statement
The information disclosure statement (IDS) submitted on 01/15/2025 has/have been considered by the examiner.
Claim Rejections - 35 USC § 112
The following is a quotation of 35 U.S.C. 112(b):
(b) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention.
The following is a quotation of 35 U.S.C. 112 (pre-AIA ), second paragraph:
The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the applicant regards as his invention.
Claims 1-8 are rejected under 35 U.S.C. 112(b) or 35 U.S.C. 112 (pre-AIA ), second paragraph, as being indefinite for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor (or for applications subject to pre-AIA 35 U.S.C. 112, the applicant), regards as the invention.
Claims 1, and 7-8 recites the limitation "a type of the modal of the input data" in lines 7, 6, and 7 respectively. There is insufficient antecedent basis for this limitation in the claim. It is not clear which modal “the modal” refers to in case of multimodal input. For examination purposes, the limitation has been interpreted as “a type of each modal of the input data”.
Claims 2-6 are also rejected under 35 U.S.C. 112(b) as being dependent upon a rejected base claim.
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claim(s) 1 and 7-8 is/are rejected under 35 U.S.C. 103 as being unpatentable over Liu et al (2021 International Conference on Neuromorphic Computing), hereinafter Liu in view of Zhou et al (arXiv:1909.11059v3 2019), hereinafter Zhou.
-Regarding claim 1, Liu discloses a learning device comprising: processing circuitry configured to (Abstract; FIGS. 1-2; TABLEs I-III; one or more processor and memory has to be used in order to implement the method shown in Zhou’s FIG. 1
PNG
media_image1.png
222
399
media_image1.png
Greyscale
): extract an encoding feature having a time series direction on a basis of input data of one or both of monomodal data that is data of a single modal or multimodal pair data including a plurality of different modals (FIG.1, Feature extraction module; Page 371, Sec. B.); embed the features of the input data (FIG. 1, Feature extraction module, Fusion module; ; Page 371, Sec. B.); connect, a plurality of segment-embedded features in the time series direction as a modal-connected feature (FIG1. Fusion module, Attention Mechanism; Page 371, Sec. B.); and calculate a model parameter using an estimated vector of a cross-modal task estimated on a basis of one or both of the segment-embedded feature or the modal-connected feature and correct data (FIG. 1, Recognition module; Page 370, Sec. C.; equations (4)-(5)).
Liu does not disclose embedding segment information that is information for identifying a type of the modal of the input data in the encoding feature on a basis of a predetermined condition. However, a person of ordinary skills in the art would understand that this is a common practice of feature embedding for multimodal (See Kitada et al (IEEE Access, 2022), hereinafter Kitada: Page 120028, 2nd Col., Sec. V.A., 2nd paragraph, “… uses token IDS (also known as segment IDs) to differentiate tokens that belong to different modalities. This is a common approach in the multimodal literature …”).
In the same field of endeavor, Zhou teaches a unified Vision-Language model for vision-language generation or understanding tasks (Zhou: Abstract; FIGS. 1-2). Zou further teaches embedding segment information that is information for identifying a type of the modal of the input data in the encoding feature on a basis of a predetermined condition (Zhou: FIG. 2, caption; Page 3, 2nd Col., 2nd paragraph).
Therefore, it would have been obvious to one of ordinary skills in the art before the effective filing date of the claimed invention to combine the teaching of Liu with the teaching of Zhou by using segment information that is information for identifying a type of the modal of the input data in the encoding feature in order to provide modality-aware embedding to map each modality’s data into a shared vector space, and thus enable reasoning and operating across multiple sensory or data types while maintaining semantic alignment (Zhou: page 4, 2nd Col., Sec. “Fine-Tuning for Downstream Tasks”, 1st paragraph).
-Regarding claim 2, Liu in view of Zhou teaches the learning device of claim 1. The combination further teaches extracting the single encoding feature in a case where the input data is the single monomodal data, and to extract the encoding feature according to a number of types of the modals included in the input data in a case where the input data is one or both of two or more of the monomodal data or the multimodal pair data (Liu: FIG. 1).
-Regarding claim 3, Liu in view of Zhou teaches the learning device of claim 2. The combination further teaches extracting the encoding feature on a basis of a neural network corresponding to the type of the modal (Liu: FIG. 1).
-Regarding claim 4, Liu in view of Zhou teaches the learning device of claim 1.The combination further teaches to embed, in the encoding feature, a vector having a same sequence length as the encoding feature as an input and including a fixed value different for each modal (Liu: FIG. 1; Page 371, Sec. B., “512 dimensional …”; Zhou: FIG. 2).
-Regarding claim 5, Liu in view of Zhou teaches the learning device of claim 1.The combination further teaches in a case of having a plurality of the segment-embedded features as inputs, the connection unit connects connect the plurality of the segment-embedded features in the time series direction (Liu: FIG. 1; Page 368, 2nd Col., 2nd paragraph, “feature-level fusion is multi-modal representation learning … learn a joint representation from a shared hidden layer connected to multiple modal inputs … model-level fusion learns the multi-modal interaction within the model …”; Page 369, 1st Col., 1st paragraph, “… based on the attention mechanism. Combines the connection between the various modals, and weakens the error caused by the differences between the various modals”; Page 370, 1st Col., Sec. B.; equations (4)-(5)).
-Regarding claim 6, Liu in view of Zhou teaches the learning device of claim 1.The combination further teaches performing conversion using a function of an arbitrary neural network on a basis of one or both of the segment-embedded feature or the modal-connected feature, and estimate a vector corresponding to the correct data as the estimated vector of the cross-modal task (Liu: FIG. 1; equations (4)-(5)).
-Regarding claim 7, Liu discloses a learning method comprising (Abstract; FIGS. 1-2; TABLEs I-III): extracting an encoding feature having a time series direction on a basis of input data of one or both of monomodal data that is data of a single modal or multimodal pair data including a plurality of different modals (FIG.1, Feature extraction module; Page 371, Sec. B.); embedding the features of the input data (FIG. 1, Feature extraction module, Fusion module; ; Page 371, Sec. B.); connecting, a plurality of segment-embedded features in the time series direction as a modal-connected feature (FIG1. Fusion module, Attention Mechanism; Page 371, Sec. B.); and calculating a model parameter using an estimated vector of a cross-modal task estimated on a basis of one or both of the segment-embedded feature or the modal-connected feature and correct data (FIG. 1, Recognition module; Page 370, Sec. C.; equations (4)-(5)).
Liu does not disclose embedding segment information that is information for identifying a type of the modal of the input data in the encoding feature on a basis of a predetermined condition. However, a person of ordinary skills in the art would understand that this is a common practice of feature embedding for multimodal (See Kitada et al (IEEE Access, 2022), hereinafter Kitada: Page 120028, 2nd Col., Sec. V.A., 2nd paragraph, “… uses token IDS (also known as segment IDs) to differentiate tokens that belong to different modalities. This is a common approach in the multimodal literature …”).
In the same field of endeavor, Zhou teaches a unified Vision-Language model for vision-language generation or understanding tasks (Zhou: Abstract; FIGS. 1-2). Zou further teaches embedding segment information that is information for identifying a type of the modal of the input data in the encoding feature on a basis of a predetermined condition (Zhou: FIG. 2, caption; Page 3, 2nd Col., 2nd paragraph).
Therefore, it would have been obvious to one of ordinary skills in the art before the effective filing date of the claimed invention to combine the teaching of Liu with the teaching of Zhou by using segment information that is information for identifying a type of the modal of the input data in the encoding feature in order to provide modality-aware embedding to map each modality’s data into a shared vector space, and thus enable reasoning and operating across multiple sensory or data types while maintaining semantic alignment (Zhou: page 4, 2nd Col., Sec. “Fine-Tuning for Downstream Tasks”, 1st paragraph).
-Regarding claim 8, Liu discloses a non-transitory computer-readable recording medium storing therein a learning program that causes a computer to execute a process comprising (Abstract; FIGS. 1-2; TABLEs I-III; one or more processor and memory has to be used in order to implement the method shown in Zhou’s FIG. 1): extracting an encoding feature having a time series direction on a basis of input data of one or both of monomodal data that is data of a single modal or multimodal pair data including a plurality of different modals (FIG.1, Feature extraction module; Page 371, Sec. B.); embedding the features of the input data (FIG. 1, Feature extraction module, Fusion module; ; Page 371, Sec. B.); connecting, a plurality of segment-embedded features in the time series direction as a modal-connected feature (FIG1. Fusion module, Attention Mechanism; Page 371, Sec. B.); and calculating a model parameter using an estimated vector of a cross-modal task estimated on a basis of one or both of the segment-embedded feature or the modal-connected feature and correct data (FIG. 1, Recognition module; Page 370, Sec. C.; equations (4)-(5)).
Liu does not disclose embedding segment information that is information for identifying a type of the modal of the input data in the encoding feature on a basis of a predetermined condition. However, a person of ordinary skills in the art would understand that this is a common practice of feature embedding for multimodal (See Kitada et al (IEEE Access, 2022), hereinafter Kitada: Page 120028, 2nd Col., Sec. V.A., 2nd paragraph, “… uses token IDS (also known as segment IDs) to differentiate tokens that belong to different modalities. This is a common approach in the multimodal literature …”).
In the same field of endeavor, Zhou teaches a unified Vision-Language model for vision-language generation or understanding tasks (Zhou: Abstract; FIGS. 1-2). Zou further teaches embedding segment information that is information for identifying a type of the modal of the input data in the encoding feature on a basis of a predetermined condition (Zhou: FIG. 2, caption; Page 3, 2nd Col., 2nd paragraph).
Therefore, it would have been obvious to one of ordinary skills in the art before the effective filing date of the claimed invention to combine the teaching of Liu with the teaching of Zhou by using segment information that is information for identifying a type of the modal of the input data in the encoding feature in order to provide modality-aware embedding to map each modality’s data into a shared vector space, and thus enable reasoning and operating across multiple sensory or data types while maintaining semantic alignment (Zhou: page 4, 2nd Col., Sec. “Fine-Tuning for Downstream Tasks”, 1st paragraph).
Conclusion
Any inquiry concerning this communication or earlier communications from the examiner should be directed to XIAO LIU whose telephone number is (571)272-4539. The examiner can normally be reached Monday-Thursday and Alternate Fridays 8:30-4:30.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Jennifer Mehmood can be reached at (571) 272-2976. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/XIAO LIU/Primary Examiner, Art Unit 2664