DETAILED ACTION
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Claims 1-17 are pending.
Claim Interpretation - 35 USC § 112(f)
The following is a quotation of 35 U.S.C. 112(f):
(f) Element in Claim for a Combination. – An element in a claim for a combination may be expressed as a means or step for performing a specified function without the recital of structure, material, or acts in support thereof, and such claim shall be construed to cover the corresponding structure, material, or acts described in the specification and equivalents thereof.
The following is a quotation of pre-AIA 35 U.S.C. 112, sixth paragraph:
An element in a claim for a combination may be expressed as a means or step for performing a specified function without the recital of structure, material, or acts in support thereof, and such claim shall be construed to cover the corresponding structure, material, or acts described in the specification and equivalents thereof.
The claims in this application are given their broadest reasonable interpretation using the plain meaning of the claim language in light of the specification as it would be understood by one of ordinary skill in the art. The broadest reasonable interpretation of a claim element (also commonly referred to as a claim limitation) is limited by the description in the specification when 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is invoked.
As explained in MPEP § 2181, subsection I, claim limitations that meet the following three-prong test will be interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph:
(A) the claim limitation uses the term “means” or “step” or a term used as a substitute for “means” that is a generic placeholder (also called a nonce term or a non-structural term having no specific structural meaning) for performing the claimed function;
(B) the term “means” or “step” or the generic placeholder is modified by functional language, typically, but not always linked by the transition word “for” (e.g., “means for”) or another linking word or phrase, such as "configured to" or "so that"; and
(C) the term “means” or “step” or the generic placeholder is not modified by sufficient structure, material, or acts for performing the claimed function.
Use of the word “means” (or “step”) in a claim with functional language creates a rebuttable presumption that the claim limitation is to be treated in accordance with 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph. The presumption that the claim limitation is interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is rebutted when the claim limitation recites sufficient structure, material, or acts to entirely perform the recited function.
Absence of the word “means” (or “step”) in a claim creates a rebuttable presumption that the claim limitation is not to be treated in accordance with 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph. The presumption that the claim limitation is not interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is rebutted when the claim limitation recites function without reciting sufficient structure, material or acts to entirely perform the recited function.
Claim limitations in this application that use the word “means” (or “step”) are being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, except as otherwise indicated in an Office action. Conversely, claim limitations in this application that do not use the word “means” (or “step”) are not being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, except as otherwise indicated in an Office action.
This application includes one or more claim limitations that do not use the word “means,” but are nonetheless being interpreted under 35 U.S.C. 112(f) or preAIA 35 U.S.C. 112, sixth paragraph, because the claim limitation(s) uses a generic placeholder that is coupled with functional language without reciting sufficient structure to perform the recited function and the generic placeholder is not preceded by a structural modifier. Such claim limitation(s) is/are:
Claim 1:
“a user input information collector configured to…”;
“a sensing information analyzer configured to…”;
“a user semantic information generator configured to…”;
“a multimodal foundation model configured to…”; and
“an image-input decoder configured to…”.
Claim 2:
“a personalized semantic model configured to…”;
“a personalized semantic model manager configured to…”; and
“a personalized semantic generator configured to…”.
Because this/these claim limitation(s) is/are being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, it/they is/are being interpreted to cover the corresponding structure described in the specification as performing the claimed function, and equivalents thereof.
If applicant does not intend to have this/these limitation(s) interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, applicant may: (1) amend the claim limitation(s) to avoid it/them being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph (e.g., by reciting sufficient structure to perform the claimed function); or (2) present a sufficient showing that the claim limitation(s) recite(s) sufficient structure to perform the claimed function so as to avoid it/them being interpreted under 35 U.S.C. 112(f) or preAIA 35 U.S.C. 112, sixth paragraph.
Claim Rejections - 35 USC § 103
The following is a quotation of pre-AIA 35 U.S.C. 103(a) which forms the basis for all obviousness rejections set forth in this Office action:
(a) A patent may not be obtained though the invention is not identically disclosed or described as set forth in section 102 of this title, if the differences between the subject matter sought to be patented and the prior art are such that the subject matter as a whole would have been obvious at the time the invention was made to a person having ordinary skill in the art to which said subject matter pertains. Patentability shall not be negatived by the manner in which the invention was made.
Claim(s) 1-2, 6-8, 10 and 14-16 is/are rejected under 35 U.S.C. 103 as being unpatentable over Lee et al (US20210390700A1) in view of Golard et al (US10909405B1).
Regarding claims 1 and 10, Lee teaches a personalized image segmentation device comprising: at least one processor; and a memory having instructions stored thereon, which, when executed by the at least one processor, cause the at least one processor to implement:
(Lee, Fig. 4; "a referring image segmentation network”, [0087])
a user input information collector configured to output input information in a second format based on a user input in a first format received from a first external device;
(Lee, Fig. 4; "language feature extractor 405 may receive audio information associated with the input image, and may identify the referral expression based on the audio information." [0062]; "Given an input image and a natural language expression from the user 100, the described referring image segmentation network (which may be implemented on server 110) may segment out an object referred by the natural language expression within the input image. User 100 may communicate with the server 110 via network 105." [0024]; receiving a user-side input from an external client (user 100 over network 105) in a first format (audio) and outputting/identifying the referral expression in a second format (text-based natural language expression) used downstream. The audio-to-referral-expression mapping satisfies the "first format → second format" conversion from a first external device)
a sensing information analyzer configured to output context information based on sensing information received from a second external device;
(Golard, Fig. 5/Fig. 6; "Sensor 240 may include a position sensor, an inertial measurement unit (IMU), a depth camera assembly, or any combination thereof." [c5:25-30]; "the data describing the interaction may be recorded by a device that operates in connection with AR device 504 (e.g., within an Internet of Things (IoT) system). As a specific example, a smart toaster within a same IoT system as AR device 504 may record that user 506 toasts a bagel every morning at around 6 am." [c12:60-c13:5]; "In certain embodiments, one or more ephemeral factors (e.g., factors predicted to be affecting a current state of user 506) may influence the interest segmentation. Such ephemeral factors may include, without limitation, a time of day, a recent activity of the user, and/or a current activity of user 506." [c13:25-35]; Golard teaches an analyzer that derives context (ephemeral factors, current activity, location/IMU) from sensing information received from a second external device (an IoT-paired device such as a smart toaster, or position/IMU sensors))
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention was made to incorporate Golard's sensing-derived context into Lee in order to make Lee's referring-image-segmentation query-context aware, improving segmentation accuracy when the user is in different real-world contexts (Lee itself recognizes the need for accurate visual-lingual alignments [0018-0019, 0094]). The combination of Lee and Gupta also teaches other enhanced capabilities.
The combination of Lee and Gupta further teaches:
a user semantic information generator configured to generate, for image segmentation, personalized user input information that identifies a target object to be segmented in image data by converting the input information into more specific input information based on the context information and the personalized semantic information associated with a user;
(Lee, Fig. 4; "language features may be extracted from a referral expression (e.g., “A girl in a brown outfit holding a game controller”) by language feature extractor 405, while image features are extracted from a corresponding image using image feature extractor 410." [0056]; Golard, Fig. 5; "performing an interest segmentation of the image to determine a personal interest that the user may have in a particular object identified via the semantic segmentation" [c1:40-45]; "Segmentation module 510 may predict user 506's personal interest in particular object 513 (that is, may perform the interest segmentation) based on one or more of a variety of factors, or on any relevant combination of factors." [c11:45-55]; "In one example, a personal profile module 516 may have created a personal profile 518 for user 506. Personal profile 518 may be generated based on a variety of factors, such as the factors described above. This personal profile may then be used to perform the interest segmentation." [c13:55-65]; Lee teaches generating input information (referral expression → language features) for segmentation, but does not personalize/contextualize it; Golard teaches refining the identification of the target object based on personalized semantic information (personal profile 518) and context (ephemeral factors). Together, Lee and Golard teach generating personalized, more-specific input information identifying the target object)
a multimodal foundation model configured to encode the image data and the personalized user input information respectively to generate feature information corresponding to the image data; and
(Lee, Fig. 4; "Embodiments of the apparatus and method may include a language feature extractor configured to extract an image feature vector from an input image, an image feature extractor configured to extract a plurality of language feature vectors for a referral expression, wherein each of the plurality of language feature vectors comprises a different number of dimensions, a fusion module configured to combine each of the language feature vectors with the image feature vector to produce a plurality of self-attention vectors and to combine the plurality of self-attention vectors to produce a multi-modal feature vector" [0006]; "embodiments of the present disclosure take an input image and a language expression and extracts visual and lingual features from each of them. Both the features are used to construct a joint multi-modal feature representation." [0026]; a multimodal model respectively encodes image data and the language/user-input information into a joint multi-modal feature representation)
an image-input decoder configured to detect the target object in the image data based on the feature information and the personalized user input information."
(Lee, Fig. 4; "a decoder configured to decode the multi-modal feature vector to produce an image mask indicating a portion of the input image corresponding to the referral expression." [0006]; "Decoder 420 decodes the multi-modal feature vector to produce an image mask indicating a portion of the input image corresponding to the referral expression." [0080]; "the system may produce a pixel-level segmentation mask of an area within an input image given a natural language expression (e.g., of the white chair in the image of the room)." [0050]; a decoder detects/segments the target object in the image based on the combined feature information and the user-input (referral expression). The decoder's image-mask output identifies the target-object pixels, => "detect the target object.")
Regarding claim 2, the combination of Lee and Golard teaches its/their respective base claim(s).
The combination further teaches the personalized image segmentation device of claim 1, wherein the user semantic information generator includes:
a personalized semantic model configured to store the personalized semantic information;
a personalized semantic model manager configured to update the personalized semantic model; and
a personalized semantic generator configured to generate the personalized user input information by converting the input information into the more specific input information based on the personalized semantic information and the context information.
(Golard, Fig. 5; "In one example, a personal profile module 516 may have created a personal profile 518 for user 506. Personal profile 518 may be generated based on a variety of factors, such as the factors described above. This personal profile may then be used to perform the interest segmentation. In some embodiments, personal profile 518 may be generated and/or maintained by a backend server. In other embodiments, personal profile 518 may be generated and/or maintained by AR device 504." [c13:55-65]; "Segmentation module 510 may use the factors described above to perform the interest segmentation in a variety of ways. In some examples, segmentation module 510 may perform the interest segmentation using a neural network." [c13:45-50]; "In some embodiments, a personalized profile may be created for the user based on the inputs that the user has opted-in to providing." [c17:40-45]; Golard teaches: (i) personal profile 518 = personalized semantic model (stores personalized info); (ii) personal profile module 516 that creates/maintains (i.e., updates) the profile = personalized semantic model manager; (iii) segmentation module 510 that uses the profile together with contextual factors to identify the personally-significant target object = personalized semantic generator. using Golard's personal-profile architecture would disambiguate Lee's referral expression among similar objects (per Lee [0019]))
Regarding claims 6 and 14, the combination of Lee and Golard teaches its/their respective base claim(s).
The combination further teaches the personalized image segmentation device of claim 2, wherein the personalized semantic model manager performs user profiling at preset periods and updates the personalized semantic information based on the user profiling.
(Golard, Fig. 5; "In one example, a personal profile module 516 may have created a personal profile 518 for user 506. Personal profile 518 may be generated based on a variety of factors, such as the factors described above. This personal profile may then be used to perform the interest segmentation." [c13:55-65]; "the user may be digitally transmitted a periodic summary of data collected relating to him or her (e.g., a daily summary, a weekly summary, a monthly summary, etc.). Specific examples of such data may include eye-tracking data, browsing history, and/or GPS data collected during the period (e.g., collected by the user's AR device and/or by an additional device such as a smartphone or a laptop)." [c16:50-60]; "the user may grant permission to allow the data to ... (2) be used to perform an operation (such as creating a personalized profile)" [c16:55-65]; Golard teaches periodic (daily/weekly/monthly) data collection and use of that data to create/maintain the personalized profile. Under KSR, performing the profile creation/update at the disclosed preset periods (daily/weekly/monthly) is the predictable application of a known technique. Golard supplies the preset-period profiling and profile update)
Regarding claims 7 and 15, the combination of Lee and Golard teaches its/their respective base claim(s).
The combination further teaches the personalized image segmentation device of claim 1, wherein the sensing information includes at least one of location information and inertial information.
(Golard, Fig. 2; "Additionally or alternatively, the interest segmentation may be based on GPS data associated with the user, URL browsing data associated with the user, and/or user-submitted data submitted by the user." [c2:10-15]; "Sensor 240 may include a position sensor, an inertial measurement unit (IMU), a depth camera assembly, or any combination thereof." [c5:25-30]; "Examples of sensor 240 may include, without limitation, accelerometers, gyroscopes, magnetometers, other suitable types of sensors that detect motion, sensors used for error correction of the IMU, or some combination thereof." [c5:30-40]; Golard teaches both alternatives: location information (GPS data, position sensor) and inertial information (IMU, accelerometers, gyroscopes))
Regarding claims 8 and 16, the combination of Lee and Golard teaches its/their respective base claim(s).
The combination further teaches the personalized image segmentation device of claim 1, wherein
the user input collector includes a speech-to-text converter, and
the first format includes a speech format and the second format includes a text format.
(Lee, "language feature extractor 405 may receive audio information associated with the input image, and may identify the referral expression based on the audio information.", [0062]; converting audio/speech information into a text-based referral expression)
Claim(s) 3, 5, 11 and 13 is/are rejected under 35 U.S.C. 103 as being unpatentable over Lee et al (US20210390700A1) in view of Golard et al (US10909405B1) and further in view of Gupta et al (US20250014287A1).
Regarding claims 3 and 11, the combination of Lee and Golard teaches its/their respective base claim(s).
The combination does not expressly disclose but Gupta teaches the personalized image segmentation device of claim 2, wherein the personalized semantic model manager determines a similarity between current input information corresponding to a current user input and previous input information corresponding to a previous user input, and when the similarity is greater than a preset threshold, updates the personalized semantic information based on the current input information.
(Gupta, "The application server 106 may then compare the computed context vector with the plurality of context vectors of the plurality of multimedia for a second level of filtering. The comparison of the computed context vector with the plurality of context vectors of the plurality of multimedia may include implementation of the various data comparison algorithms, for example, approximate nearest neighbor algorithm, cosine similarity algorithm, or the like.", [0057]; "The application server 106 may then select at least one of the plurality of multimedia content for which contextual relevance with the first target object 104 or the second user of the second mobile device exceeds a threshold value", [0059]; computing similarity between vectors and using a threshold value to update/select context info)
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention was made to incorporate the vector similarity comparison and threshold-based filtering of Gupta into the personalized image segmentation system of Lee and Golard in order to systematically evaluate incoming user inputs before modifying the user's personalized semantic profile. The combination of Lee and Gupta also teaches other enhanced capabilities.
Regarding claims 5 and 13, the combination of Lee, Golard and Gupta teaches its/their respective base claim(s).
The combination further teaches the personalized image segmentation device of claim 3, wherein the similarity is determined based on word similarity and contextual similarity between the input information and the previous input information.
(Gupta, "The context vector may correspond to a numerical representation (for example, an 'n' dimensional array) of a context of a multimedia content. The application server 106 may utilize one or more known techniques (for example, word-to-vector technique) to generate a plurality of context vectors", [0042]; determining context vectors and similarities utilizing word-to-vector similarity techniques)
Claim(s) 4 and 12 is/are rejected under 35 U.S.C. 103 as being unpatentable over Lee et al (US20210390700A1) in view of Golard et al (US10909405B1) and further in view of Gupta et al (US20250014287A1) and Ramanujapuram et al (US20090285492A1).
Regarding claims 4 and 12, the combination of Lee, Golard and Gupta teaches its/their respective base claim(s).
The combination does not expressly disclose but Ramanujapuram teaches the personalized image segmentation device of claim 3, wherein the personalized semantic model manager, when the similarity is equal to or less than the threshold, generates the personalized semantic information based on the input information.
(Ramanujapuram, "the server may check whether a previously stored histogram (or salient points, etc.) matches a histogram (or salient points, etc.) of the newly received image within a predefined matching threshold.", [0053]; "the server stores location information, histogram information, OCR information, operation results, or other data. The stored information is generally indexed to the captured image, so that the stored information can be used to evaluate a subsequent captured image.", [0064]; checking a predefined matching threshold, and when an input's similarity fails to meet the threshold, it extracts and stores new image and contextual parameters to evaluate subsequent inputs, functionally generating new personalized semantic information)
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention was made to incorporate the teachings of Ramanujapuram into the modified system or method of Lee and Golard in order to continuously update and adapt the system's knowledge base when encountering novel user inputs. The combination of Lee, Gupta and Ramanujapuram also teaches other enhanced capabilities.
Claim(s) 9 and 17 is/are rejected under 35 U.S.C. 103 as being unpatentable over Lee et al (US20210390700A1) in view of Golard et al (US10909405B1) and further in view of Hwang et al (US20130108115A1).
Regarding claims 9 and 17, the combination of Lee and Golard teaches its/their respective base claim(s).
The combination does not expressly disclose but Hwang teaches the personalized image segmentation device of claim 1, wherein
the user input collector includes a handwriting-to-text converter, and
the first format includes an image format and the second format includes a text format.
(Hwang, "Optical character recognition (OCR) is a mechanical or electronic translation of scanned images of handwritten, typewritten or printed text, graphics or symbols into machine-encoded text.", [0002]; utilizing an OCR system as a handwriting-to-text converter, where the first format is scanned images and the second format is machine-encoded text.)
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention was made to incorporate the teachings of Hwang into the modified system or method of Lee and Golard in order to expand the user input collector to support natural handwriting inputs captured via images. The combination of Lee, Golard and Hwang also teaches other enhanced capabilities.
Response to Arguments
Applicant's arguments filed on 6/9/2026 with respect to one or more of the pending claims which were rejected under 35 USC § 103 have been fully considered but are moot in view of the new ground(s) of rejection.
Applicant also arguments that the claim interpretation under 35 USC § 112(f) invoked in the previous office action should be overcome by the amendment to claims 1 and 2.
Examiner respectfully disagrees.
The amendments to claims 1 and 2 do not overcome the 112(f) interpretation because they still fail to provide specific physical structure. Adding a generic "processor and memory" to the beginning of Claim 1 is not enough; under patent rules, general-purpose computer hardware does not give structure to software claims. Furthermore, the generic placeholder words identified by the Examiner, such as "collector," "analyzer," "generator," "model," and "decoder", are left completely unchanged. To overcome a 112(f) issue, an applicant must add details about how a component is built, such as specific hardware circuits or a detailed step-by-step algorithm. However, the recent amendments only add more functional language explaining what the components do. For instance, stating that a generator will "generate... personalized user input information... by converting the input information" simply describes a task and its end result. Because the claims still use generic placeholder terms paired only with descriptions of their functions rather than concrete structures, they remain subject to 112(f) interpretation.
Conclusion
THIS ACTION IS MADE FINAL. Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a).
A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to JIANXUN YANG whose telephone number is (571)272-9874. The examiner can normally be reached on MON-FRI: 8AM-5PM Pacific Time.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Amandeep Saini can be reached on (571)272-3382. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center. for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272- 1000.
/JIANXUN YANG/
Primary Examiner, Art Unit 2662 8/8/2026