DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Priority
This application is a continuation of US Application No. 17/586,613, filed on January 27, 2022.
Information Disclosure Statement
The information disclosure statement (IDS) submitted is considered by the examiner.
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claim(s) 1-3, 8-10 and 15-17 is/are rejected under 35 U.S.C. 103 as being unpatentable over Xia et al. (US Patent Number 11,875,250 B1, hereinafter “Xia”).
(1) regarding claim 1:
As shown in fig. 1, Xia disclosed a method (col. 4, lines 52-54, note that IG. 1 illustrates an example system environment in which a neural network model whose loss function takes semantic relationships among classes into account may be used for classification), comprising:
obtaining a digital canvas comprising a plurality of layers (col. 5, 1-3, note that an image data set 204 comprising a plurality of images 287 (e.g., 287A) is used as input), wherein a layer of the plurality of layers comprises a semantic label associated with a digital content item depicted in the layer (col. 6, lines 37-40, note that the model 202 may comprise a number of layers in the depicted embodiment, such as convolution layers C1 and C2, pooling or sub-sampling layers P1 and P2, and fully-connected layers F1 and F2. Also, see col. 8, lines 7-9, note that Sub-graph 391 may represent semantic relationship information associated with vehicles, as suggested by the label of the root node 301);
determining a semantic graph of the plurality of layers, wherein the semantic graph comprises nodes corresponding to semantic labels associated with digital content items depicted in the digital canvas and edges linking nodes, wherein the edges comprise semantic information corresponding to the linked nodes (col. 8, lines 2-14, note that FIG. 3 illustrates an example semantic graph indicating relationships between various classes for which a neural network based classifier may be used, according to at least some embodiments. In the depicted embodiment, a semantic graph may comprise at least three sub-graphs 391, 392 and 393. Sub-graph 391 may represent semantic relationship information associated with vehicles, as suggested by the label of the root node 301. Sub-graph 392 may represent semantic relationship information associated with household items as indicated by the label of node 351, while sub-graph 393 may encode semantic relationship information associated with buildings as indicated by the label for root node 371).
Xia disclosed most of the subject matter as described as above except for specifically teaching generating, using the semantic graph, a placement context associated with the digital content item depicted in the layer of the digital canvas, wherein the placement context corresponds to a relationship between the digital content item depicted in the layer of the digital canvas and digital content items depicted in one or more other layers.
However, it would have been obvious for Xia to teach generating, using the semantic graph, a placement context associated with the digital content item depicted in the layer of the digital canvas, wherein the placement context corresponds to a relationship between the digital content item depicted in the layer of the digital canvas and digital content items depicted in one or more other layers (col. 9, lines 58-60, note that a tree generator 418 may combine linked pairs of words/terms of the sparse knowledge graph to generate a set of trees. If, for example, a seed word W1 is linked to a child word W2 in the sparse knowledge graph, and W2 itself is linked to another word W3, a tree in which W1 is the parent of W2, and W2 is the parent of W3 may be generated i.e. relation. Also see, col. 10, lines 22-29, the strengths of the relationships between pairs of classes included in a knowledge graph may be encoded as, or represented by, a square semantic distance matrix with each of the classes represented by a respective row and column, and a numerical score at element [j, k] of the matrix indicating the strength of the semantic relationship between the word/term corresponding to the jth and kth rows/columns).
At the time of filing for the invention, it would have been obvious to a person of ordinary skilled in the art for Xia to teach generating, using the semantic graph, a placement context associated with the digital content item depicted in the layer of the digital canvas, wherein the placement context corresponds to a relationship between the digital content item depicted in the layer of the digital canvas and digital content items depicted in one or more other layers. The suggestion/motivation for doing so would have been in order to obtain semantic relationships among class (abs.). Therefore, it would have been obvious for Xia to obtain the invention as specified in claim 1.
(2) regarding claim 2:
Xia further disclosed the method of claim 1, wherein each layer of the plurality of layers organizes digital content items depicted in the digital canvas based on a semantic knowledge of each digital content item depicted the layer and a semantic knowledge of the digital content items depicted in one or more other layers of the plurality of layers of the digital canvas (col. 7, lines 43-54, note that using a knowledge graph or some other source of semantic relationship information, a semantic distance matrix 249 may be obtained in the depicted embodiment, which indicates the relative similarities and/or differences between pairs of the classes. The information in the semantic distance matrix and the difference vector may be aggregated or combined in one embodiment, e.g., by multiplying the difference vector with the distance matrix, to obtain a semantically-weighted loss vector 252. The semantically-weighted loss vector may then be used to modify parameters such as weights and biases at various layers of the model 202 in the depicted embodiment).
(3) regarding claim 3:
Xia further disclosed the method of claim 1, wherein the relationship of the digital content item depicted in the layer of the digital canvas is a spatial relationship and a semantic relationship of the digital content items depicted in one or more other layers of the plurality of layers of the digital canvas (col. 11, lines 37-43, note that the improved grouping and sub-grouping enabled by the use of semantic relationship information may help provide better search responses in some embodiments. For example, in one embodiment, an image management service at which semantic relationship information is used during training of a neural network may receive a search request for an object similar to a specified object. Also note that instead of responding with an exact match for a search query, a search result may provide examples of related objects (objects which belong to semantically similar classes, not necessarily the exact class of the searched-for object) in some embodiments using the semantic information incorporated in the feature vectors produced at one or more layers of the neural network.).
The proposed rejection of claims 1-3 renders obvious the non-transitory computer readable medium claims (823, fig. 8) 8-10 and the system claims (figs. 1 and 2) 15-17 because these steps occur in the operation of the proposed rejection as discussed above. Thus, the arguments similar to that presented above for claims 1-3 are equally applicable to claims 8-10 and 15-17.
Claim(s) 4-7, 11-14 and 18-20 is/are rejected under 35 U.S.C. 103 as being unpatentable over Xia et al. (US Patent Number 11,875,250 B1, hereinafter “Xia”) in view of Burachas et al. (US Publication Number 2019/0370587 A1, hereinafter “Burachas”).
(1) regarding claim 4:
Xia further disclosed the method of claim 1, further comprising: comparing the placement context to a content rule (col. 11, lines 44-49, note that using the trained neural network model with the semantic information incorporated within it, the service may generate a feature vector corresponding to the image depicting the Siamese cat, and compare that feature vector to feature vectors associated with previously-analyzed images);
adjusting the digital content item according to the content rule (col. 11, lines 54-59, note that instead of responding with an exact match for a search query, a search result may provide examples of related objects (objects which belong to semantically similar classes) i.e. adjusting the rule for cats).
Xia disclosed most of the subject matter as described as above except for specifically teaching updating an edge linking a first node of the semantic graph associated with the digital content item and a second node of the semantic graph associated with a class of the digital content item.
However, Burachas disclosed updating an edge linking a first node of the semantic graph associated with the digital content item and a second node of the semantic graph associated with a class of the digital content item (para. [0058], note that VQA model 20 may update the tree data structure to create a link (or, in other words, edge) between the root node and each sub-area segmented from the area including the person, creating child nodes for each sub-area including the shirt, head, pant or short, and any other semantically relevant area of the area including the person, and specifying in each child node the corresponding label of “shirt,” “head,” “short” or “pant,” etc. Also see, VQA model 20 may label each edge with a relationship between the root node and the child node. For example, VQA model 20 may label the edge between the root “person” node and the child “shirt” node).
At the time of filing for the invention, it would have been obvious to a person of ordinary skilled in the art to teach updating an edge linking a first node of the semantic graph associated with the digital content item and a second node of the semantic graph associated with a class of the digital content item. The suggestion/motivation for doing so would have been in order to obtain an artificial intelligence model may, when analyzing the image to output the result, segment the image into hierarchically arranged semantic areas in which objects in the image are segmented into parts, determine, based on the query, an attention mask for the areas, update, based on the attention mask, the image to visually identify which of the areas formed a basis for the result, and output the updated image (abs.). Therefore, it would have been obvious to combine Xia with Burachas to obtain the invention as specified in claim 4.
(2) regarding claim 5:
Xia disclosed most of the subject matter as described as above except for specifically teaching wherein the content rule defines a class of the digital content items depicted in the digital canvas.
However, Burachas disclosed wherein the content rule defines a class of the digital content items depicted in the digital canvas (para. [0079], note that the hierarchical stages analyze content of the images (or videos) and identify semantic areas (or, in other words, regions), such as sub-parts, parts, objects, attributes, and relationships between object pairs and entire scene types (and also potentially elementary actions and complex actions for video).).
At the time of filing for the invention, it would have been obvious to a person of ordinary skilled in the art to teach wherein the content rule defines a class of the digital content items depicted in the digital canvas. The suggestion/motivation for doing so would have been in order to obtain an artificial intelligence model may, when analyzing the image to output the result, segment the image into hierarchically arranged semantic areas in which objects in the image are segmented into parts, determine, based on the query, an attention mask for the areas, update, based on the attention mask, the image to visually identify which of the areas formed a basis for the result, and output the updated image (abs.). Therefore, it would have been obvious to combine Xia with Burachas to obtain the invention as specified in claim 5.
(3) regarding claim 6:
Xia further disclosed the method of claim 1, further comprising:
comparing the placement context to a content rule associated with a second digital content item depicted in the digital canvas (col. 11, lines 44-49, note that using the trained neural network model with the semantic information incorporated within it, the service may generate a feature vector corresponding to the image depicting the Siamese cat, and compare that feature vector to feature vectors associated with previously-analyzed images);
adjusting the digital content item according to the content rule (col. 11, lines 54-59, note that even if the image management service does not find another image with a Siamese cat, the fact that the class “Siamese cat” is semantically close to other cat-related classes such as “Persian cat” may enable examples of the similar classes to be found (e.g., based on similar feature vectors) and returned in response to the search request for images of Siamese cats).
Xia disclosed most of the subject matter as described as above except for specifically teaching updating an edge linking a first node of the semantic graph associated with the digital content item and a second node of the semantic graph associated with the second digital content item.
However, Burachas disclosed updating an edge linking a first node of the semantic graph associated with the digital content item and a second node of the semantic graph associated with the second digital content item (para. [0058], note that VQA model 20 may update the tree data structure to create a link (or, in other words, edge) between the root node and each sub-area segmented from the area including the person, creating child nodes for each sub-area including the shirt, head, pant or short, and any other semantically relevant area of the area including the person, and specifying in each child node the corresponding label of “shirt,” “head,” “short” or “pant,” etc. Also see, VQA model 20 may label each edge with a relationship between the root node and the child node. For example, VQA model 20 may label the edge between the root “person” node and the child “shirt” node).
At the time of filing for the invention, it would have been obvious to a person of ordinary skilled in the art to teach updating an edge linking a first node of the semantic graph associated with the digital content item and a second node of the semantic graph associated with the second digital content item. The suggestion/motivation for doing so would have been in order to obtain an artificial intelligence model may, when analyzing the image to output the result, segment the image into hierarchically arranged semantic areas in which objects in the image are segmented into parts, determine, based on the query, an attention mask for the areas, update, based on the attention mask, the image to visually identify which of the areas formed a basis for the result, and output the updated image (abs.). Therefore, it would have been obvious to combine Xia with Burachas to obtain the invention as specified in claim 6.
(4) regarding claim 7:
Xia disclosed most of the subject matter as described as above except for specifically teaching wherein adjusting the digital content item according to the content rule further comprises: moving the digital content item a number of pixels, or moving a z-order of the digital content item.
However, Burachas disclosed wherein adjusting the digital content item according to the content rule further comprises: moving the digital content item a number of pixels, or moving a z-order of the digital content item (para. [0065], note that Scene-graph generation module 110 may identify the one or more objects, parts, sub-parts, etc. at a pixel-level of granularity (meaning, for example, that each pixel representative of an object, a part, a sub-part, etc. is individually associated with the object, part, sub-part, etc.). As such, scene-graph generation module 110 may segment the one or more objects into one or more parts at the pixel-level of granularity).
At the time of filing for the invention, it would have been obvious to a person of ordinary skilled in the art to teach wherein adjusting the digital content item according to the content rule further comprises: moving the digital content item a number of pixels, or moving a z-order of the digital content item. The suggestion/motivation for doing so would have been in order to obtain an artificial intelligence model may, when analyzing the image to output the result, segment the image into hierarchically arranged semantic areas in which objects in the image are segmented into parts, determine, based on the query, an attention mask for the areas, update, based on the attention mask, the image to visually identify which of the areas formed a basis for the result, and output the updated image (abs.). Therefore, it would have been obvious to combine Xia with Burachas to obtain the invention as specified in claim 7.
The proposed rejection of claims 4-7 renders obvious the non-transitory computer readable medium claims (823, fig. 8) 11-14 and the system claims 18-20 (figs. 1 and 2) because these steps occur in the operation of the proposed rejection as discussed above. Thus, the arguments similar to that presented above for claims 4-7 are equally applicable to claims 11-14 and 18-20.
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure.
Wang et al. (US Patent Number 11,113,599 B2) disclosed methods and systems for generating captions for digital images. In particular, the disclosed systems and methods can train an image encoder neural network and a sentence decoder neural network to generate a caption from an input digital image.
Any inquiry concerning this communication or earlier communication from the examiner should be directed to Hilina K Demeter whose telephone number is (571) 270-1676.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, King Y. Poon could be reached at (571) 270- 0728. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of an application may be obtained from the Patent Application Information Retrieval (PAIR) system. Status information for published applications may be obtained from either Private PAIR or Public PAIR. Status information for unpublished applications is available through Private PAIR only. For more information about PAIR system, see http://pari-direct.uspto.gov. Should you have questions on access to the Private PAIR system, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative or access to the automated information system, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/HILINA K DEMETER/Primary Examiner, Art Unit 2617