Prosecution Insights
Last updated: August 17, 2026
Application No. 19/312,499

STORING ENTRIES IN AND RETRIEVING INFORMATION FROM AN EMBEDDING OBJECT MEMORY

Non-Final OA §101§103
Filed
Aug 28, 2025
Priority
Dec 19, 2022 — provisional 63/433,619 +1 more
Examiner
HOANG, SON T
Art Unit
Tech Center
Assignee
Microsoft Technology Licensing, LLC
OA Round
1 (Non-Final)
84%
Grant Probability
Favorable
1-2
OA Rounds
1y 11m
Est. Remaining
99%
With Interview

Examiner Intelligence

Grants 84% — above average
84%
Career Allowance Rate
769 granted / 920 resolved
+23.6% vs TC avg
Strong +35% interview lift
Without
With
+34.7%
Interview Lift
resolved cases with interview
Typical timeline
2y 11m
Avg Prosecution
10 currently pending
Career history
932
Total Applications
across all art units

Statute-Specific Performance

§101
16.1%
-23.9% vs TC avg
§103
54.8%
+14.8% vs TC avg
§102
12.2%
-27.8% vs TC avg
§112
6.1%
-33.9% vs TC avg
Black line = Tech Center average estimate • Based on career data from 920 resolved cases

Office Action

§101 §103
Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Status This instant application No. 19/312,499 has claims 21-40 pending based on the preliminary amendment filed on September 2, 2025. Priority / Filing Date Applicant’s claim for priorities of parent application No. 18/122,563 (now Pat. No. US 12405934) and provisional application No. 63/433,619 is acknowledged. The effective filing date for this application is December 19, 2022. Abstract The abstract of the disclosure is objected due to the use of implied language. Note that in the abstract, the language should be clear and concise and should not repeat information given in the title. It should avoid using phrases which can be implied, such as, “The disclosure concerns,” “The disclosure defined by this invention,” “The disclosure describes,” etc… See MPEP § 608.01(b). Note that in the abstract, Applicant cites “Methods, systems, and media for storing entries in and/or retrieving information from an embedding object memory are provided” on lines 1-2. This citation clearly provokes the use of implied language. Correction is required (e.g., removal of the entire first sentence of the abstract). Drawings The drawings filed on August 28, 2025 are acceptable for examination purposes. Information Disclosure Statement As required by M.P.E.P. 609(C), the Applicant’s submissions of the Information Disclosure Statements filed on 4 September 2025, 29 October 2025, 30 January 2026, and 23 April 2026 are acknowledged by the Examiner, and the cited references have been considered in the examination of the claims now pending. As required by M.P.E.P. 609 C(2), a copy of each of the PTOL-1449s initialed and dated by the Examiner is attached to the instant Office action. Double Patenting The nonstatutory double patenting rejection is based on a judicially created doctrine grounded in public policy (a policy reflected in the statute) so as to prevent the unjustified or improper timewise extension of the “right to exclude” granted by a patent and to prevent possible harassment by multiple assignees. A nonstatutory double patenting rejection is appropriate where the claims at issue are not identical, but at least one examined application claim is not patentably distinct from the reference claim(s) because the examined application claim is either anticipated by, or would have been obvious over, the reference claim(s). See, e.g., In re Berg, 140 F.3d 1428, 46 USPQ2d 1226 (Fed. Cir. 1998); In re Goodman, 11 F.3d 1046, 29 USPQ2d 2010 (Fed. Cir. 1993); In re Longi, 759 F.2d 887, 225 USPQ 645 (Fed. Cir. 1985); In re Van Ornum, 686 F.2d 937, 214 USPQ 761 (CCPA 1982); In re Vogel, 422 F.2d 438, 164 USPQ 619 (CCPA 1970); and In re Thorington, 418 F.2d 528, 163 USPQ 644 (CCPA 1969). A timely filed terminal disclaimer in compliance with 37 CFR 1.321(c) or 1.321(d) may be used to overcome an actual or provisional rejection based on a nonstatutory double patenting ground provided the reference application or patent either is shown to be commonly owned with this application, or claims an invention made as a result of activities undertaken within the scope of a joint research agreement. See MPEP § 717.02 for applications subject to examination under the first inventor to file provisions of the AIA as explained in MPEP § 2159. See MPEP §§ 706.02(l)(1) - 706.02(l)(3) for applications not subject to examination under the first inventor to file provisions of the AIA . A terminal disclaimer must be signed in compliance with 37 CFR 1.321(b). The USPTO Internet website contains terminal disclaimer forms which may be used. Please visit www.uspto.gov/forms/. The filing date of the application in which the form is filed determines what form (e.g., PTO/SB/25, PTO/SB/26, PTO/AIA /25, or PTO/AIA /26) should be used. A web-based eTerminal Disclaimer may be filled out completely online using web-screens. An eTerminal Disclaimer that meets all requirements is auto processed and approved immediately upon submission. For more information about eTerminal Disclaimers, refer to http://www.uspto.gov/patents/process/file/efs/guidance/eTD-info-I.jsp. Claims 21-40 are rejected on the ground of nonstatutory double patenting over claims 1-20 of Pat. No. US 12405934. Claims 21-40 of the instant application recite similar limitations and claims 1-20 of ‘934 as being compared in the table below. For the purpose of illustration, only claims 21-27, and 35-40 (method claims) of the instant application are compared to the claims of the patent (underlining are used to indicate conflict limitations). The remaining claims of the instant application recite different categories (i.e., system claims) and are therefore not compared for simplicity purposes. Instant Application Pat. No. US 12405934 Claim 21 A method for storing entries in an embedding object memory, the method comprising: receiving a plurality of content items, the plurality of content items including a first content item of a first content type and a second content item of a second content type, the second content type being different than the first content type; providing the first content item of the first content type to a first semantic embedding model trained to generate one or more semantic embeddings; receiving, from the first semantic embedding model, one or more first semantic embeddings; providing the second content item of the second content type to a second semantic embedding model, wherein the second semantic embedding model is also trained to generate one or more semantic embeddings; receiving, from the second semantic embedding model, one or more second semantic embeddings; inserting the one or more first semantic embeddings corresponding to the first content item of the first content type and the one or more second semantic embeddings corresponding to the second content item of the second content type into the embedding object memory, the insertion triggering a spatial storage operation to store one or more vector representations of the one or more first semantic embeddings and the one or more second semantic embeddings; and providing the embedding object memory. Claim 1 A method for storing an entry in an embedding object memory, the method comprising: receiving a content item, the content item having one or more content data; providing one of the content data associated with the content item to one or more semantic embedding models, wherein the one or more semantic embedding models generate one or more semantic embeddings; receiving, from one or more of the semantic embedding models, one or more semantic embeddings, wherein a collection of semantic embeddings is associated with a first semantic embedding model of the one or more semantic embedding models, wherein the collection of semantic embeddings comprises a first semantic embedding generated by the first semantic embedding model for at least one content data from the respective content item, wherein the one or more semantic embedding models comprise a version, and wherein each of the semantic embeddings generated by each of the respective one or more semantic embedding models comprise metadata corresponding to the version; inserting the one or more semantic embeddings into the embedding object memory, wherein the embedding object memory stores one or more semantic embeddings from the collection of semantic embeddings, wherein the one or more semantic embeddings are associated with a respective indication corresponding to a reference to source data associated with the one or more semantic embeddings, and wherein the insertion triggers a spatial storage operation to store a vector representation of the one or more semantic embeddings; providing an updated semantic embedding model to replace at least one of the semantic embedding models… providing the embedding object memory. See further Divakaran and Oleson below for mapping details and motivation to combine with the claims of ‘934. Claim 22 The method of claim 21, wherein the first content type is image data, and wherein the first semantic embedding model is trained to generate one or more semantic embeddings from image input. Claim 9 The method of claim 8, wherein the source data includes one or more of audio files, text files, or image files. Claim 1 providing one of the content data associated with the content item to one or more semantic embedding models, wherein the one or more semantic embedding models generate one or more semantic embeddings… Claim 23 The method of claim 22, wherein the second content type is text data, and wherein the second semantic embedding model is trained to generate one or more semantic embeddings from text input. Claim 9 The method of claim 8, wherein the source data includes one or more of audio files, text files, or image files. Claim 1 providing one of the content data associated with the content item to one or more semantic embedding models, wherein the one or more semantic embedding models generate one or more semantic embeddings… Claim 24 The method of claim 22, wherein the second content type is audio data, and wherein the second semantic embedding model is trained to generate one or more semantic embeddings from audio input. Claim 9 The method of claim 8, wherein the source data includes one or more of audio files, text files, or image files. Claim 1 providing one of the content data associated with the content item to one or more semantic embedding models, wherein the one or more semantic embedding models generate one or more semantic embeddings… Claim 25 The method of claim 21, wherein the plurality of content items further include a third content item of a third content type, the third content type being different than the first and second content types, and the method further comprising: providing the third content item of the third content type to a third semantic embedding model trained to generate one or more semantic embeddings; receiving, from the third semantic embedding model, one or more third semantic embeddings; and inserting the one or more third semantic embeddings corresponding to the third content item of the third content type into the embedding object memory. Claim 9 The method of claim 8, wherein the source data includes one or more of audio files, text files, or image files. Claim 1 providing one of the content data associated with the content item to one or more semantic embedding models, wherein the one or more semantic embedding models generate one or more semantic embeddings… receiving, from one or more of the semantic embedding models, one or more semantic embeddings… inserting the one or more semantic embeddings into the embedding object memory, wherein the embedding object memory stores one or more semantic embeddings from the collection of semantic embeddings… Claim 26 The method of claim 25, wherein the third content type is a skill. Claim 1 in view of Oleson. See mapping and motivation to combine below. Claim 27 The method of claim 21, wherein the one or more vector representations of the one or more first semantic embeddings and the one or more second semantic embeddings are stored in at least one of an approximate nearest neighbor (ANN) tree, a k-d tree, or a multidimensional tree. Claim 2 The method of claim 1, wherein the vector representation is stored in at least one of an approximate nearest neighbor (ANN) tree, a k-d tree, or a multidimensional tree. Claim 35 A method for storing entries in an embedding object memory, the method comprising :receiving a plurality of content items, the plurality of content items including a first content item of a first content type, a second content item of a second content type, and a third content item of a third content type, the first content type, the second content type, and the third content type being different than each other; generating one or more first semantic embeddings based on the first content item of the first content type, using a first semantic embedding model; inserting the one or more first semantic embeddings corresponding to the first content item of the first content type into the embedding object memory, the insertion triggering a first spatial storage operation to store one or more first vector representations of the one or more first semantic embeddings; generating one or more second semantic embeddings based on the second content item of the second content type, using a second semantic embedding model; inserting the one or more second semantic embeddings corresponding to the second content item of the second content type into the embedding object memory, the insertion triggering a second spatial storage operation to store one or more second vector representations of the one or more second semantic embeddings; generating one or more third semantic embeddings based on the third content item of the third content type, using a third semantic embedding model; inserting the one or more third semantic embeddings corresponding to the third content item of the third content type into the embedding object memory, the insertion triggering a third spatial storage operation to store one or more third vector representations of the one or more third semantic embeddings; and providing the embedding object memory comprising the one or more first vector representations of the one or more first semantic embeddings corresponding to the first content item of the first content type, the one or more second vector representations of the one or more second semantic embeddings corresponding to the second content item of the second content type, and the one or more third vector representations of the one or more third semantic embeddings corresponding to the third content item of the third content type. Claim 1 A method for storing an entry in an embedding object memory, the method comprising: receiving a content item, the content item having one or more content data; providing one of the content data associated with the content item to one or more semantic embedding models, wherein the one or more semantic embedding models generate one or more semantic embeddings; receiving, from one or more of the semantic embedding models, one or more semantic embeddings, wherein a collection of semantic embeddings is associated with a first semantic embedding model of the one or more semantic embedding models, wherein the collection of semantic embeddings comprises a first semantic embedding generated by the first semantic embedding model for at least one content data from the respective content item…; inserting the one or more semantic embeddings into the embedding object memory, wherein the embedding object memory stores one or more semantic embeddings from the collection of semantic embeddings, wherein the one or more semantic embeddings are associated with a respective indication corresponding to a reference to source data associated with the one or more semantic embeddings, and wherein the insertion triggers a spatial storage operation to store a vector representation of the one or more semantic embeddings; providing an updated semantic embedding model to replace at least one of the semantic embedding models… receiving, from the updated semantic embedding model, an updated one or more semantic embeddings corresponding to the one or more semantic embeddings generated by the at least one of the semantic embedding models, wherein the updated one or more semantic embeddings are generated based on the one of the content data used to generate the one or more semantic embeddings; inserting the updated semantic embeddings in the embedding object memory with metadata corresponding to the updated version; and providing the embedding object memory. Claim 9 The method of claim 8, wherein the source data includes one or more of audio files, text files, or image files. See further Divakaran and Oleson below for mapping details and motivation to combine with the claims of ‘934. Claim 36 The method of claim 35, wherein the first content type is image data, and wherein the first semantic embedding model is trained to generate one or more semantic embeddings from image input. Claim 9 The method of claim 8, wherein the source data includes one or more of audio files, text files, or image files. Claim 37 The method of claim 36, wherein the second content type is audio data, and wherein the second semantic embedding model is trained to generate one or more semantic embeddings from audio input. Claim 9 The method of claim 8, wherein the source data includes one or more of audio files, text files, or image files. Claim 38 The method of claim 36, wherein the second content type is text data, and wherein the second semantic embedding model is trained to generate one or more semantic embeddings from text input. Claim 9 The method of claim 8, wherein the source data includes one or more of audio files, text files, or image files. Claim 39 The method of claim 38, wherein the third content type is a skill. Claim 1 in view of Oleson. See mapping and motivation to combine below. Claim 40 The method of claim 39, wherein the one or more first vector representations, the one or more second vector representations, and the one or more third vector representations are stored in at least one of an approximate nearest neighbor (ANN) tree, a k-d tree, or a multidimensional tree. Claim 2 The method of claim 1, wherein the vector representation is stored in at least one of an approximate nearest neighbor (ANN) tree, a k-d tree, or a multidimensional tree. Although the conflicting claims are not identical, they are not patentably distinct from each other because they are substantially similar in scope, and they use the similar limitations to produce the same end result of storing entries in and retrieving information from an embedding object memory. It would have been obvious to a person with ordinary skills in the art at the time of the invention was effectively filed to modify the elements of claims 1-20 of ‘934 with any combination of the cited references below to arrive at the pending claims of the instant application for the purpose of supporting content recommendations and/or user preference determination of multimodal content items based on matching features and representations within an embedding vector space to improve similarity-based retrieval efficiency and organization using multimodal embeddings. Further, it would have been obvious to a person with ordinary skills in the art at the time of the invention was effectively filed to modify or to omit the additional elements of claims 1-20 of ‘934 to arrive at the pending claims of the instant application because the person would have realized that the remaining element would perform the same functions as before. “Omission of element and its function in combination is obvious expedient if the remaining elements perform same functions as before.” See In re Karlson (CCPA) 136 USPQ 184, decide Jan 16, 1963, Appl. No. 6857, U.S. Court of Customs and Patent Appeals. Claim Rejections - 35 USC § 101 35 U.S.C. 101 reads as follows: Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title. The claimed invention in claims 21-40 is directed to a judicial exception (i.e., an abstract idea) without significantly more. Claims 21-40 pass step 1 of the 35 U.S.C. 101 analysis since each claim is either directed to a method, a system comprising a processor, and memory (i.e., hardware components per [0141] of instant specification). a. Claim 21 recites, in part, steps that are directed to an abstract idea (“Courts have examined claims that required the use of a computer and still found that the underlying, patent-ineligible invention could be performed via pen and paper or in a person’s mind.” Versata Dev. Group v. SAP Am., Inc., 793 F.3d 1306, 1335, 115 USPQ2d 1681, 1702 (Fed. Cir. 2015)) per step 2A – prong 1 since the core of each claim recites steps or functions of receiving content items of different types, providing them to trained semantic embedding models, receiving semantic embeddings, inserting them into an embedding object memory to trigger a spatial storage operation to store vector representations, and providing the memory. These limitations fall within the recognized judicial categories of: Mathematical Concepts/Algorithms: generating and storing semantic embeddings and vector representations reply on mathematical functions and multi-dimensional spatial calculations Certain Methods of Organizing Human Activity/Mental Processes: classifying and organizing data items by content type into a structured memory system. Per step 2A - prong 2, the claim does not integrate the abstract idea into a practical application. The claim elements perform generic data gathering (“receiving”), model invocation (“providing”) , vector extraction (“receiving”), and data storage (“inserting…triggering a spatial storage operation”, “providing”). These limitations do not impose any meaningful technical constraints on how the system executes the spatial storage operation. The claim merely recites the result of storing vectors upon insertion rather than a specific structural implementation or concrete algorithm that improves the operational hardware of the computer system. Per step 2B, considering the claims elements individually and as an ordered combination, the hardware and functional execution components are described at a high level of generality (e.g., standard servers, microprocessors, standard neural net embedding models). Executing conventional data manipulation steps (e.g., extracting embeddings and storing vector representations in memory) using generic computing infrastructure represents well-understood, routine, and conventional (WURC) activity in the art. The ordered combination of steps merely amounts to an instruction to apply the abstract mathematical and data organization concepts using standard computer components. b. Claim 28 recites system hardware (“a processor; and memory”) executing functional instructions identical in operational scope to claim 21. Per step 2A - prong 2, reciting a generic processor and memory does not integrate the abstract concept of multi-model embedding generation and spatial storage into a practical application. The claim fails to receive technical limitations detailing how the processor or memory architecture is physically altered, reconfigured, or optimized to handle spatial storage beyond standard execution of software routines. Per step 2B, the physical hardware components processor and memory are generic off-the-shelf components performing their expected basic computing duties. The combination of generic server hardware executing high-level software operations, e.g., receiving multi-model data, invoking models, generating and storing vectors, does not supply significantly more than the abstract idea itself. c. Claim 35 explicitly enumerates generating first, second, and third embeddings for three distinct content types and triggering respective first, second, and third spatial storage operations. Per step 2A - prong 2, sequentially repeating the abstract process across three data types does not alter the underlying nature of the functional operations. The claim recites the result of triggering sequential storage operations without defining any concrete, non-conventional technical mechanism or specialized database architecture that achieves an improvement in memory throughput or retrieval latency. Per step 2B, simply repeating generic data-processing steps three times for three different data types adds no technical substance or inventive concept to the ordered combination. The limitation remains a high-level operational concept executed on WURC computing devices. b. Claims 22-24, 26, 29-31, 33, and 36-39 further recite data type and input modality limitations. Per step 2A – prong 2, these claims restrict the input content items to specific data types (e.g., image data, text data, audio data, or a skill) and specify that the corresponding models are trained on those inputs. Restricting an abstract idea to a specific technical field or data format (e.g., images or audio) does not integrate the judicial exception into a practical application per MPEP 2106.05(h). Per step 2B, specifying that the data manipulated consists of audio, image, text, or skills relies on WURC data modalities handled by standard ML models. Restrictions to field-of-use do not provide an inventive concept or amount to significantly more. c. Claims 25, and 32 recite the multimodal pipeline requiring that the third content type processed wherein the third content type is different from the first and second type. Per step 2A – prong 2, the addition of a third content item, a third semantic embedding model, and a third semantic embedding merely expands the volume and types of data processed by the underlying algorithm. Limiting an abstract process to additional data inputs or specific data categories does not integrate the judicial exception into a practical application. Per step 2B, the claims do not recite elements sufficient to transform the abstract idea into a patent-eligible application. The recited hardware components are invokes performing their WURC functions in the art. Re-executing the generic steps of providing, receiving, and inserting for a third semantic embedding adds no physical technical improvement to the execution of the computer hardware or memory and fails to supply an inventive concept. d. Claims 27, 34, and 40 recite specific spatial data structures that the vector representations are stored in. Per step 2A – prong 2, while these claims recite specify concrete database structures, they merely state that the vectors are stored in these conventional indexing structures. Without claiming specific structural modifications to these trees or technical rules for operating them to resolve specific hardware bottleneck, the claims fail to integrate the abstract idea into a practical application. Per step 2B, storage using ANN trees, k-d trees, or multidimensional trees is WURC techniques used in the spatial database and vector retrieval arts. Reciting conventional spatial index structures without technical improvements to their underlying operation fails to supply an inventive concept. Claim Rejections - 35 USC § 103 The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claims 21-25, 27-32, 34-38, and 40 are rejected under 35 U.S.C. 103 as being unpatentable over Divakaran et al. (Pub. No. US 2021/0297498, published on September 23, 2021; hereinafter Divakaran) in view of Yin et al. (Pub. No. US 2024/0289313, filed on December 14, 2022; hereinafter Yin). Regarding claims 21, and 28, Divakaran clearly shows and discloses a method for storing entries in an embedding object memory (Abstract); and a system for storing entries in an embedding object memory, the system comprising: a processor; and memory storing instructions that, when executed by the processor, cause the system to perform a set of operations of the method (Figure 10 and texts), the method comprising: receiving a plurality of content items, the plurality of content items including a first content item of a first content type and a second content item of a second content type, the second content type being different than the first content type (multimodal content (e.g., video images, audio, text and the like) associated with a user(s) and information regarding a user(s) can be monitored and separated by the optional CCM module 110, into different modalities, such as, text, images, audio, user information and the like, and can be communicated to an appropriate embedding module. That is, in some embodiments the optional CCM module 110 can monitor a user's interaction with content associated with, for example, a social media site, [0033]); providing the first content item of the first content type to a first semantic embedding model (Image portions of the content can be communicated to the image embedding module 130 for embedding into a common embedding space as described below, [0033]) trained to generate one or more semantic embeddings (the image embedding module 130 determines respective vector representations of the images separated from content associated with the user(s) for embedding the images into the common embedding space 210. The image embedding module 130 can include a deep image encoder 220 implementing a convolutional neural network (CNN) with fully connected (FC) layers to process the images for embedding, [0040]); receiving, from the first semantic embedding model, one or more first semantic embeddings (second modality feature vector representation of image content of the multimodal content can be created by the image embedding module 130 for embedding into the common embedding space, [0042]); providing the second content item of the second content type to a second semantic embedding model (text portions of the content can be communicated to the text embedding module 120 for embedding into the common embedding space as described below, [0033]), wherein the second semantic embedding model is also trained to generate one or more semantic embeddings (the text embedding module 120 can include a deep text encoder 230 implementing a text-based convolutional neural network (CNN) with fully connected (FC) layers to process the text for embedding, [0040]); receiving, from the second semantic embedding model, one or more second semantic embeddings (first modality vector representation of word/text content of the multimodal content can be created by the text embedding module 120 for embedding into a common embedding space, [0042]); inserting the one or more first semantic embeddings corresponding to the first content item of the first content type and the one or more second semantic embeddings corresponding to the second content item of the second content type into the embedding object memory (a semantic embedding space can be trained using a plurality of text and images (and audio, etc.) having known categories. Vector representations of the text/images and respective categories of the text/images are embedded in the semantic embedding space such that related text/images and categories of the content that are closer in the embedding space than unrelated text/images and categories, [0074]); and providing the embedding object memory (using a predefined semantic embedding space as described above, the generator 152 can determine features of the determined user preferred audio content that define the determined user preferred audio content as comprising an English accent. A second audio content can include a message/intent desired to be conveyed using the determined user preferred content type (e.g., English accent). For example, the second audio content can include a message/intent in an Irish accent conveying a message that music is therapeutic. In such an embodiment, the discriminator 154 can determine the intent of the second audio content using a predetermined semantic embedding space as described above, [0075]). Yin then discloses the insertion triggering a spatial storage operation to store one or more vector representations of the one or more first semantic embeddings and the one or more second semantic embeddings (Data points 404 in dataset 402 include various types and/or formats of data. For example, data points 404 could include images, text, audio, video, point clouds, meshes, time series data, and/or other types of data in a high-dimensional space, [0059]. Processing engine 322 organizes embeddings 408 into a tree structure 406 that includes a hierarchy of nodes 414(1)-414(3) (each of which is referred to individually as node 414). In one or more embodiments, tree structure 406 includes a ball tree, KD-tree, or another type of tree that is used to spatially organize embeddings 408 into disjoint subsets, [0062], [0094]. It is clear that inserting embeddings dynamically builds or partitions nodes in a spatial tree structure, e.g., creating root/child nodes defined by hypersphere radiuses and centroids). It would have been obvious to an ordinary person skilled in the art at the time of the invention was effectively filed to incorporate the teachings of Yin with the teachings of Divakaran for the purpose of supporting content recommendations and/or user preference determination of multimodal content items based on matching features and representations within an embedding vector space to improve similarity-based retrieval efficiency and organization using multimodal embeddings. Regarding claim 35, Divakaran clearly shows and discloses a method for storing entries in an embedding object memory (Abstract), the method comprising: receiving a plurality of content items, the plurality of content items including a first content item of a first content type, a second content item of a second content type, and a third content item of a third content type, the first content type, the second content type, and the third content type being different than each other (multimodal content (e.g., video images, audio, text and the like) associated with a user(s) and information regarding a user(s) can be monitored and separated by the optional CCM module 110, into different modalities, such as, text, images, audio, user information and the like, and can be communicated to an appropriate embedding module. That is, in some embodiments the optional CCM module 110 can monitor a user's interaction with content associated with, for example, a social media site, [0033]); generating one or more first semantic embeddings based on the first content item of the first content type (Image portions of the content can be communicated to the image embedding module 130 for embedding into a common embedding space as described below, [0033]), using a first semantic embedding model (the image embedding module 130 determines respective vector representations of the images separated from content associated with the user(s) for embedding the images into the common embedding space 210. The image embedding module 130 can include a deep image encoder 220 implementing a convolutional neural network (CNN) with fully connected (FC) layers to process the images for embedding, [0040]); inserting the one or more first semantic embeddings corresponding to the first content item of the first content type into the embedding object memory (a semantic embedding space can be trained using a plurality of text and images (and audio, etc.) having known categories. Vector representations of the text/images and respective categories of the text/images are embedded in the semantic embedding space such that related text/images and categories of the content that are closer in the embedding space than unrelated text/images and categories, [0074]); generating one or more second semantic embeddings based on the second content item of the second content type (text portions of the content can be communicated to the text embedding module 120 for embedding into the common embedding space as described below, [0033]), using a second semantic embedding model (the text embedding module 120 can include a deep text encoder 230 implementing a text-based convolutional neural network (CNN) with fully connected (FC) layers to process the text for embedding, [0040]); inserting the one or more second semantic embeddings corresponding to the second content item of the second content type into the embedding object memory (a semantic embedding space can be trained using a plurality of text and images (and audio, etc.) having known categories. Vector representations of the text/images and respective categories of the text/images are embedded in the semantic embedding space such that related text/images and categories of the content that are closer in the embedding space than unrelated text/images and categories, [0074]); generating one or more third semantic embeddings based on the third content item of the third content type (audio portions of the content can be communicated to the optional audio embedding module 170 for converting to text and embedding into the common embedding space as described below, [0033]), using a third semantic embedding model (a respective modality vector representation of audio content of the multimodal content can be created by the image audio embedding module 170 for embedding into the common embedding space, [0042]); inserting the one or more third semantic embeddings corresponding to the third content item of the third content type into the embedding object memory (a semantic embedding space can be trained using a plurality of text and images (and audio, etc.) having known categories. Vector representations of the text/images and respective categories of the text/images are embedded in the semantic embedding space such that related text/images and categories of the content that are closer in the embedding space than unrelated text/images and categories, [0074]); and providing the embedding object memory comprising the one or more first vector representations of the one or more first semantic embeddings corresponding to the first content item of the first content type, the one or more second vector representations of the one or more second semantic embeddings corresponding to the second content item of the second content type, and the one or more third vector representations of the one or more third semantic embeddings corresponding to the third content item of the third content type (using a predefined semantic embedding space as described above, the generator 152 can determine features of the determined user preferred audio content that define the determined user preferred audio content as comprising an English accent. A second audio content can include a message/intent desired to be conveyed using the determined user preferred content type (e.g., English accent). For example, the second audio content can include a message/intent in an Irish accent conveying a message that music is therapeutic. In such an embodiment, the discriminator 154 can determine the intent of the second audio content using a predetermined semantic embedding space as described above, [0075]). Yin then discloses: the insertion triggering a first spatial storage operation to store one or more first vector representations of the one or more first semantic embeddings; the insertion triggering a second spatial storage operation to store one or more second vector representations of the one or more second semantic embeddings; the insertion triggering a third spatial storage operation to store one or more third vector representations of the one or more third semantic embeddings (Data points 404 in dataset 402 include various types and/or formats of data. For example, data points 404 could include images, text, audio, video, point clouds, meshes, time series data, and/or other types of data in a high-dimensional space, [0059]. Processing engine 322 organizes embeddings 408 into a tree structure 406 that includes a hierarchy of nodes 414(1)-414(3) (each of which is referred to individually as node 414). In one or more embodiments, tree structure 406 includes a ball tree, KD-tree, or another type of tree that is used to spatially organize embeddings 408 into disjoint subsets, [0062], [0094]. It is clear that inserting embeddings dynamically builds or partitions nodes in a spatial tree structure, e.g., creating root/child nodes defined by hypersphere radiuses and centroids). It would have been obvious to an ordinary person skilled in the art at the time of the invention was effectively filed to incorporate the teachings of Yin with the teachings of Divakaran for the purpose of supporting content recommendations and/or user preference determination of multimodal content items based on matching features and representations within an embedding vector space to improve similarity-based retrieval efficiency and organization using multimodal embeddings. Regarding claims 22, 29, and 36, Divakaran further discloses the first content type is image data (Image portions of the content can be communicated to the image embedding module 130 for embedding into a common embedding space as described below, [0033]), and wherein the first semantic embedding model is trained to generate one or more semantic embeddings from image input (the image embedding module 130 determines respective vector representations of the images separated from content associated with the user(s) for embedding the images into the common embedding space 210. The image embedding module 130 can include a deep image encoder 220 implementing a convolutional neural network (CNN) with fully connected (FC) layers to process the images for embedding, [0040]). Regarding claims 23, 30, and 38, Divakaran further discloses the second content type is text data (text portions of the content can be communicated to the text embedding module 120 for embedding into the common embedding space as described below, [0033]), and wherein the second semantic embedding model is trained to generate one or more semantic embeddings from text input (the text embedding module 120 can include a deep text encoder 230 implementing a text-based convolutional neural network (CNN) with fully connected (FC) layers to process the text for embedding, [0040]). Regarding claims 24, 31, and 37, Divakaran further discloses the second content type is audio data (audio portions of the content can be communicated to the optional audio embedding module 170 for converting to text and embedding into the common embedding space as described below, [0033]), and wherein the second semantic embedding model is trained to generate one or more semantic embeddings from audio input (a respective modality vector representation of audio content of the multimodal content can be created by the image audio embedding module 170 for embedding into the common embedding space, [0042]). Regarding claims 25, and 32, Divakaran further discloses the plurality of content items further include a third content item of a third content type, the third content type being different than the first and second content types, and the method further comprising: providing the third content item of the third content type to a third semantic embedding model trained to generate one or more semantic embeddings (audio portions of the content can be communicated to the optional audio embedding module 170 for converting to text and embedding into the common embedding space as described below, [0033]); receiving, from the third semantic embedding model, one or more third semantic embeddings (a respective modality vector representation of audio content of the multimodal content can be created by the image audio embedding module 170 for embedding into the common embedding space, [0042]); and inserting the one or more third semantic embeddings corresponding to the third content item of the third content type into the embedding object memory (a semantic embedding space can be trained using a plurality of text and images (and audio, etc.) having known categories. Vector representations of the text/images and respective categories of the text/images are embedded in the semantic embedding space such that related text/images and categories of the content that are closer in the embedding space than unrelated text/images and categories, [0074]). Regarding claims 27, 34, and 40, Yin further discloses the one or more vector representations of the one or more first semantic embeddings and the one or more second semantic embeddings are stored in at least one of an approximate nearest neighbor (ANN) tree, a k-d tree, or a multidimensional tree (Within the tree structure, a given parent node stores embeddings of a set of data points that is bounded by a hypersphere or another region of the multidimensional vector space occupied by the embeddings. Two or more child nodes of the parent node store disjoint subsets of the embeddings from the parent node, [0056]. Processing engine 322 organizes embeddings 408 into a tree structure 406 that includes a hierarchy of nodes 414(1)-414(3) (each of which is referred to individually as node 414). In one or more embodiments, tree structure 406 includes a ball tree, KD-tree, or another type of tree that is used to spatially organize embeddings 408 into disjoint subsets, [0062], [0094]). Claims 26, 33, and 39 are rejected under 35 U.S.C. 103 as being unpatentable over Divakaran in view of Yin and further in view of Oleson et al. (Pub. No. US 2024/0152546, filed on November 6, 2023; hereinafter Oleson). Regarding claims 26, 33, and 39, Oleson then discloses the third content type is a skill (The image of the diagram can then be processed by the one or more diagram analysis models 120 (e.g., a multimodal embedding model and/or a diagram parsing model) to generate relevant search results for the diagram in the image. The multimodal embedding model can generate, in some embodiments, a text search query and/or embeddings of the text and images in the diagram. Relevant search results from a multimodal embedding model can include horizontal search features, such as skills, concepts, practice problems, relevant videos, equations, and the like, as well as similar images for identifying similar diagrams to the input diagram, [0035]. It is clear that the multimodal embedding model generates embeddings for these features to retrieve relevant results wherein “skill” is mapped into an embedding space). It would have been obvious to an ordinary person skilled in the art at the time of the invention was effectively filed to incorporate the teachings of Oleson with the teachings of Divakaran, as modified by Yin, for the purpose of enabling unified cross-modal semantic retrieval within a single spatially indexed memory structure. Relevant Prior Art The following references are deemed relevant to the claims: Lin et al. (Pub. No. US 2021/0271707) teaches generating a search result based on receiving a query having a text input by a joint embedding model trained to generate an image result. Training the joint embedding model includes accessing a set of images and textual information and generating a set of image-text pairs based on matches between image feature vectors and textual feature vectors. Additionally, operations include generating an image result for display by the joint embedding model based on the text input.. Newman et al. (Pat. No. US 11106736) teaches compiling, using at least one processor, a corpus of training data by obtaining, for each respective concept object in an ontology, a respective concept label and respective annotations describing the respective concept object; generating a vocabulary of terms based on the corpus of training data; training a semantic model using the corpus of training data and the vocabulary of terms, wherein input features for the semantic model are based on context words in proximity to a term in the vocabulary of terms; and storing a set of word embeddings for the vocabulary of terms based on the trained semantic model. Contact Information Any inquiry concerning this communication or earlier communications from the Examiner should be directed to Son Hoang whose telephone number is (571) 270-1752. The Examiner can normally be reached on Monday – Friday (7:00 AM – 4:00 PM). If attempts to reach the Examiner by telephone are unsuccessful, the Examiner’s supervisor, Sherief Badawi can be reached on (571) 272-9782. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of an application may be obtained from the Patent Application Information Retrieval (PAIR) system. Status information for published applications may be obtained from either Private PAIR or Public PAIR. Status information for unpublished applications is available through Private PAIR only. For more information about the PAIR system, see http://pair-direct.uspto.gov. Should you have questions on access to the Private PAIR system, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative or access to the automated information system, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /SON T HOANG/Primary Examiner, Art Unit 2169 July 24, 2026
Read full office action

Prosecution Timeline

Aug 28, 2025
Application Filed
Jul 28, 2026
Non-Final Rejection mailed — §101, §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12688256
CENTRALIZED REPOSITORY AND DATA SHARING HUB FOR ESTABLISHING MODEL SUFFICIENCY
4y 2m to grant Granted Jul 21, 2026
Patent 12688229
MEDIA FILE RECOMMENDATIONS FOR A SEARCH ENGINE
1y 6m to grant Granted Jul 21, 2026
Patent 12664210
MEDIA FILE RECOMMENDATIONS FOR A SEARCH ENGINE
1y 5m to grant Granted Jun 23, 2026
Patent 12639308
QUERY PROCESSING DEVICE AND QUERY PROCESSING METHOD
1y 7m to grant Granted May 26, 2026
Patent 12632476
APPARATUSES, METHODS, AND COMPUTER PROGRAM PRODUCTS FOR PROVIDING PREDICTIVE INFERENCES RELATED TO A GRAPH REPRESENTATION OF DATA VIA AN APPLICATION PROGRAMMING INTERFACE
2y 4m to grant Granted May 19, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
84%
Grant Probability
99%
With Interview (+34.7%)
2y 11m (~1y 11m remaining)
Median Time to Grant
Low
PTA Risk
Based on 920 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month