DETAILED ACTION
This Office action is in response to a communication filed on June 17th, 2026, for Application No. 18/145,301, in which claims 1-19 are presented for examination. The amendments filed on June 17th, 2026, have been entered, where claims 1 and 9-10 are amended.
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Information Disclosure Statement
The information disclosure statements submitted on 12/22/2022, 04/11/2024, 06/19/2025, 11/07/2025, 05/11/2026, and 06/18/2026 are in compliance with the provisions of 37 CFR 1.97. Accordingly, the information disclosure statements were considered by the examiner.
Specification
The contents of the specification are sufficient for examination purposes.
Claim Rejections - 35 USC § 101
35 U.S.C. 101 reads as follows:
Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title.
Claims 18-19 are rejected under 35 U.S.C. 101 because the claimed invention is directed to abstract ideas without significantly more.
Regarding Claim 18:
Step 1: Claim 18 is a method claim. Therefore, claims 18-19 are directed to a statutory category of eligible subject matter.
Step 2A Prong 1: If a claim limitation, under its broadest reasonable interpretation, covers performance of the limitation in the mind but for the recitation of generic computer components, then it falls within the "Mental Processes" grouping of abstract ideas. Here, steps of the claimed method are mental processes. Specifically, the claim recites
“A method for visual content processing, comprising: . . . a plurality of predictions for the media content” (mental process – amounts to exercising judgment to evaluate known or observed media content in order to form opinions on predictions, which may be aided by pen and paper) and
“selecting a subset of the media content based on the plurality of predictions output” (mental process – amounts to exercising judgment to form an opinion on a subset of known or observed media content to select, with reference to known or observed predictions, which may be aided by pen and paper).
Step 2A Prong 2: This judicial exception is not integrated into a practical application. The claim recites the additional elements:
“applying a first machine learning model to media content . . . wherein the first machine learning model outputs . . . by the first machine learning model . . . to a second machine learning model” (amounts to mere instructions to apply the judicial exception on generic and unspecialized computer components, which do not impose any meaningful limits on practicing the abstract idea);
“wherein the first machine learning model is produced by training a student model using outputs of a teacher model . . . wherein a domain used by the first machine learning model is smaller than a domain used by the second machine learning model” (amounts to merely generally linking the use of the judicial exception to a particular technological environment or field of use, which do not impose any meaningful limits on practicing the abstract idea); and
“providing the selected subset of media content as inputs” (amounts to insignificant
extra-solution activity, merely providing the subset as input is incidental to the functioning of the claimed method).
Step 2B: The claim does not include additional elements considered individually and in combination that are sufficient to amount to significantly more than the judicial exception.
“applying a first machine learning model to media content . . . wherein the first machine learning model outputs . . . by the first machine learning model . . . to a second machine learning model” (mere instructions to apply the exception using generic computer components does not provide an inventive concept);
“wherein the first machine learning model is produced by training a student model using outputs of a teacher model . . . wherein a domain used by the first machine learning model is smaller than a domain used by the second machine learning model” (merely generally linking the use of the judicial exception to a particular technological environment or field of use does not provide an inventive concept); and
“providing the selected subset of media content as inputs” (receiving and transmitting data, such as through a network (see buySAFE, Inc. v. Google, Inc., 765 F.3d 1350, 1355, 112 USPQ2d 1093, 1096 (Fed. Cir. 2014)) or accessing information in memory (see Versata Dev. Group, Inc. v. SAP Am., Inc., 793 F.3d 1306, 1334, 115 USPQ2d 1681, 1701 (Fed. Cir. 2015); OIP Techs., 788 F.3d at 1363, 115 USPQ2d at 1092-93), is well‐understood, routine, and conventional; which is recited here with a high level of generality, and remains insignificant extra-solution activity even upon reconsideration).
For the reasons above, Claim 18 is rejected as being directed to an abstract idea without significantly more. This rejection applies equally to dependent claim 19. The additional limitations of the dependent claim are addressed below.
Regarding Claim 19:
Step 2A Prong 1: See the rejection of Claim 18 above, which Claim 19 depends on. Here, the claim recites additional elements that are mental processes. Specifically, the claim recites
“cropping from among the plurality of images based on the plurality of predictions . . . in order to create a plurality of cropped images” (mental process – apart from the “cropping” itself, which may require a particular technological environment, amounts to exercising judgement to form an opinion on known or observed data, with reference to known or observed predictions, which may be aided by pen and paper).
Step 2A Prong 2: This judicial exception is not integrated into a practical application.
The claim recites the additional element:
“wherein the media content includes a plurality of images, further comprising: cropping . . . wherein the selected subset of the media content includes the plurality of cropped images” (amounts to merely generally linking the use of the judicial exception to a particular technological environment or field of use, which do not impose any meaningful limits on practicing the abstract idea) and
“by the first machine learning model” (amounts to mere instructions to apply the judicial exception on generic and unspecialized computer components, which do not impose any meaningful limits on practicing the abstract idea).
Step 2B: The claim does not include additional elements considered individually and in combination that are sufficient to amount to significantly more than the judicial exception.
“wherein the media content includes a plurality of images, further comprising: cropping . . . wherein the selected subset of the media content includes the plurality of cropped images” (merely generally linking the use of the judicial exception to a particular technological environment or field of use does not provide an inventive concept) and
“by the first machine learning model” (mere instructions to apply the exception using generic computer components does not provide an inventive concept).
Accordingly, Claim 19 is rejected as being directed to an abstract idea without significantly more.
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claims 1, 5-6, 8-10, 14-15, and 17-19 are rejected under 35 U.S.C. 103 as being unpatentable over Arcadu et al. (hereinafter Arcadu) (Pat. App. Pub. No. US 2024/0338826 A1) in view of Chen et al. (hereinafter Chen) (“Learning Efficient Object Detection Models with Knowledge Distillation”).
Regarding Claim 1, Arcadu teaches a method for visual content processing, comprising (Abstract, “methods disclosed herein relate generally to systems and methods for detection, segmentation and characterization of isolated or overlapping object instances in digital images”, where the “methods” are for visual content processing, “for detection, segmentation and characterization . . . in digital images”; see also Fig. 3):
obtaining a subset of media content selected based on outputs of a first machine learning model (Fig. 3 and Para. [0045], “The trained DNN1 can be fed with the input image 302 to generate a region proposal 304A within the image where the at least one object of interest 302A can be located . . . At 306, a second deep neural network DNN2 can be executed on the at least one Region of Interest (RoI) 304A defined at 304”, where a subset of media content, “a region proposal 304A within the image”, which must be obtained for “a second deep neural network DNN2” to “be executed on” it, is selected based on outputs of a first machine learning model, “The trained DNN1 can . . . generate a region proposal 304A”),
wherein the first machine learning model is produced by training . . . (Para. [0045], “The trained DNN1”; see also Para. [0016], “he following described implementations refer to aspects of the disclosed systems and methods employed to detect and segment one or more isolated or overlapping objects of interest in an image by application of one or more trained Artificial Neural Networks (ANNs) or supervised Neural Networks, whereby the networks are trained on images containing one or more objects of interest with accompanying labelling metadata”),
wherein the outputs of the first machine learning model include a plurality of first predictions for a plurality of portions of the media content (Fig. 3 and Para. [0045], “The trained DNN1 can be able to generate several region proposals within the image, one for each object of interest, if several objects of interest, isolated or overlapping, are present in the input image”, where the “several region proposals” are a plurality of first predictions for a plurality of portions of the media content),
wherein the outputs of the first machine learning model are generated by applying the first machine learning model to a set of media content (Fig. 3 and Para. [0044] - [0045], “The input image 302 can comprise at least one object of interest 302A . . . At 304, a first deep neural network DNN1 can be executed. The trained DNN1 can be fed with the input image 302 to generate a region proposal 304A within the image where the at least one object of interest 302A can be located. The trained DNN1 can be able to generate several region proposals within the image”, where the outputs, “several region proposals within the image”, of the first machine learning model, “The trained DNN1”, are generated by applying the first machine learning model to a set of media content, “The trained DNN1 can be fed with the input image 302 to generate a region proposal 304A”; see also Para. [0043], “FIG. 3 illustrates an exemplary workflow for the operation of a system for object detection, segmentation and characterization in an image using for example four trained deep neural networks”),
wherein the subset of media content is a subset of the set of media content (Fig. 3 and Para. [0045], “The trained DNN1 can be able to generate several region proposals within the image, one for each object of interest, if several objects of interest, isolated or overlapping, are present in the input image. In some embodiments, the DNN1 architecture . . . for feature map extraction from the input image . . . and generat[ion of] the Region of Interest (RoI)”, where the subset of media content, “several region proposals”, is a subset of the set of media content, “within the image . . . present in the input image . . . from the input image”); and
applying a second machine learning model to the obtained subset of media content (Fig. 3 and Para. [0045], “At 306, a second deep neural network DNN2 can be executed on the at least one Region of Interest (RoI) 304A defined at 304”, where a second machine learning model, “a second deep neural network DNN2”, is applied to the obtained subset of media content, “executed on the at least one Region of Interest (RoI) 304A defined at 304”),
wherein the second machine learning model outputs a plurality of second predictions for respective portions of the plurality of portions (Fig. 3 and Para. [0045], “The trained DNN1 can be able to generate several region proposals within the image, one for each object of interest, if several objects of interest, isolated or overlapping, are present in the input image . . . The trained DNN2 can perform the classification of the one or more objects of interest in each Region of Interest (RoI), and can return a class 306A for the object of interest”, where, when “several objects of interest . . . are present in the input image”, the second machine learning model, “The trained DNN2”, outputs a plurality of second predictions, “return a class 306A for the object of interest”, for respective portions of the plurality of portions, “objects of interest in each Region of Interest (RoI)”),
. . . wherein the domain used by each machine learning model is a set of features recognized by the machine learning model . . . (Para. [0034], “the trained DNN can be a trained Convolutional Neural Network (CNN). Processing units in the early layers of CNNs learn to activate in response to simple local features, for example patterns at particular orientations or edges, while units in the deeper layers combine the low-level features into more complex patterns”; see also Fig. 3 and Para. [0045], “At 304, a first deep neural network DNN1 can be executed. The trained DNN1 can be fed with the input image 302 . . . At 306, a second deep neural network DNN2 can be executed”).
Arcadu does not explicitly disclose . . . a student model using outputs of a teacher model . . . wherein a domain used by the first machine learning model is a subset of a domain used by the second machine learning model . . . wherein the domain used by the first machine learning model is a smaller set of features than the domain used by the second machine learning model.
However, Chen teaches . . . [a method with a first and second machine learning model] (Pg. 3, Para. 3, “Method . . . we adopt the Faster-RCNN [32] as the object detection framework. Faster-RCNN
is composed of three modules: 1) A shared feature extraction through convolutional layers, 2) a
region proposal network (RPN) that generates object proposals, and 3) a classification and regression
network (RCN) that returns the detection score as well as a spatial adjustment vector for each object
proposal”, where “a region proposal network (RPN)” is comparable to Arcadu’s first machine learning model and “a classification and regression network (RCN)” is comparable to Arcadu’s second machine learning model)
[wherein the first machine learning model is produced by training] a student model using outputs of a teacher model . . . (Pg. 3, Para. 4, “We learn strong but efficient student object detectors by using the knowledge of a high capacity teacher detection network for all the three components”; see also Pg. 4, Para. 2, “Let t be the teacher model . . . The student s is trained to optimize the following loss function . . . It is known that a deep teacher can better fit to the training data and perform better in test scenarios. The soft labels contain information about the relationship between different classes as discovered by teacher. By learning from soft labels, the student network inherits such hidden information”)
wherein a domain used by the first machine learning model is a subset of a domain used by the second machine learning model . . . wherein the domain used by the first machine learning model is a smaller set of features than the domain used by the second machine learning model (Pg. 3, Para. 3, “we adopt the Faster-RCNN [32] as the object detection framework. Faster-RCNN is composed of three modules: 1) A shared feature extraction through convolutional layers, 2) a region proposal network (RPN) that generates object proposals, and 3) a classification and regression network (RCN) that returns the detection score as well as a spatial adjustment vector for each object proposal. Both the RCN and RPN use the output of 1) as features, RCN also takes the result of RPN as input”, where a domain used by the first machine learning model, “RPN use[s] the output of 1) as features”, is a subset of a domain used by the second machine learning model, “RCN . . . use[s] the output of 1) as features, RCN also takes the result of RPN as input”, such that the domain used by the first machine learning model is a smaller set of features than the domain used by the second machine learning model because it does not include “the result of RPN as input”).
Before the effective filing date of the invention, it would have been obvious to one of ordinary skill in the art to combine the method with a first and second machine learning model, wherein the first machine learning model is produced by training and wherein the domain used by each machine learning model is a set of features recognized by the machine learning model of Arcadu with the method with a first and second machine learning model, wherein the first machine learning model is produced by training a student model using outputs of a teacher model and wherein a domain used by the first machine learning model is a subset of a domain used by the second machine learning model such that the domain used by the first machine learning model is a smaller set of features than the domain used by the second machine learning model of Chen in order reduce image processing runtimes with improved accuracy using knowledge distillation and hint learning to develop a unified end-to-end multi-category framework (Chen, Pg. 1, Abstract, “Despite significant accuracy improvement in convolutional neural networks (CNN) based object detectors, they often require prohibitive runtimes to process an image . . . we propose a new framework to learn compact and fast object detection networks with improved accuracy using knowledge distillation [20] and hint learning”; see also Chen, Pg. 3, Para. 3, “we adopt the Faster-RCNN [32] as the object detection framework. Faster-RCNN is composed of three modules: 1) A shared feature extraction through convolutional layers, 2) a region proposal network (RPN) that generates object proposals, and 3) a classification and regression network (RCN) that returns the detection score as well as a spatial adjustment vector for each object proposal. Both the RCN and RPN use the output of 1) as features, RCN also takes the result of RPN as input”; see generally Chen, Pg. 2, Para. 4, “Faster-RCNN [32] and R-FCN [29], that unify various steps in object detection into an end-to-end multi-category framework”), where a teacher model can be used to learn strong models with improved generalization capacity and accuracy (Chen, Pg. 2, Para. 3, “, a teacher bounded regression loss for knowledge distillation (Section 3.3) and adaptation layers for hint learning that allows the student to better learn from the distribution of neurons in intermediate layers of the teacher” and Chen, Pg. 8, Para. 1, “In general, distillation mostly improves the generalization capability of student, while hint learning helps improving both the training and testing accuracy”; see also Chen, Pg. 3, Para. 3, “In order to achieve highly accurate object detection results, it is critical to learn strong models for all the three components”), such that the classification modules of both the first and second models can be strengthened through joint optimization (Chen, Pg. 3, Para. 4, “we learn stronger classification modules in both RPN and RCN using the knowledge distillation framework”, where “both RPN and RCN” comprise “classification modules” and Chen, Pg. 3, Fig. 1, “The two networks both use multi-task loss to jointly learn the classifier and bounding-box regressor”).
Regarding Claim 5, Arcadu in view of Chen teach the method of claim 1, further comprising: enriching the subset of media content based on the plurality of second predictions to create a set of enriched media content (Arcadu, Fig. 3 and Arcadu, Para. [0045] – [0047], “The trained DNN2 can perform the classification of the one or more objects of interest in each Region of Interest (RoI), and can return a class 306A for the object of interest . . . At 312, features 312A can be extracted from the one or more classified, detected and segmented object of interest”, where enrichment to create a set of enriched media content, “features 312A can be extracted”, is based on the plurality of second predictions, “from the one or more classified”, which, as show in figure 3, is the subset of media content, “304A”, after being “Classif[ied]”, “Detect[ed]”, and “Segment[ed]”).
Regarding Claim 6, Arcadu in view of Chen teach the method of claim 5, further comprising: sending the set of enriched media content to be used for populating a dashboard (Arcadu, Fig. 2 and Arcadu, Para [0037] – [0039], “The data processing apparatus 102 can include an Input/Output (I/O) unit 202 . . . The I/O unit 202 can comprise suitable logic, circuitry and interfaces that can act as interface between a user and the data processing apparatus 102. The I/O unit 202 can be configured to receive an input image 110 containing at least one object of interest 112. The I/O unit 202 can include different operational components of the data processing apparatus 102. The I/O unit 202 can be programmed to provide a GUI 202A for user interface . . . the GUI 202A can comprise suitable logic, circuitry and interfaces that can be configured to provide the communication between a user and the data processing apparatus 102”, where “the GUI 202A . . . can be configured to provide the communication between a user and the data processing apparatus 102” such that “The I/O unit 202 can be programmed to provide a GUI 202A for user interface”, which is within the broadest reasonable interpretation of populating a dashboard, using “output” from “the data processing apparatus 102”; see also Arcadu, Para. [0028], “The data processing device 102 can be designed to receive the input image 110 and sequentially perform detection and segmentation of the at least one object of interest 112 in the input image 110 via at least one trained Deep Neural Network (DNN)”, where “The data processing device 102” performs the operations of “receiv[ing] the input image 110 and sequentially perform[ing] detection and segmentation of the at least one object of interest 112 in the input image 110 via at least one trained Deep Neural Network (DNN)”, which as discussed above and as depicted in Fig. 3, produces the set of enriched media content, such that the output sent from the “data processing device 102” and used for populating the “GUI” dashboard is the set of enriched media content).
Regarding Claim 8, Arcadu in view of Chen teach the method of claim 1, wherein each of the first machine learning model and the second machine learning model is a classifier (Arcadu, Fig. 3 and Arcadu, Para. [0045], “At 304, a first deep neural network DNN1 can be executed. The trained DNN1 can be fed with the input image 302 to generate a region proposal 304A within the image where the at least one object of interest 302A can be located. The trained DNN1 can be able to generate several region proposals within the image, one for each object of interest, if several objects of interest, isolated or overlapping, are present in the input image. In some embodiments, the DNN1 architecture can implement algorithms like Region Proposal Networks (RPN) . . . At 306, a second deep neural network DNN2 can be executed on the at least one Region of Interest (RoI) 304A defined at 304. The trained DNN2 can perform the classification of the one or more objects of interest in each Region of Interest (RoI), and can return a class 306A for the object of interest”, where the first machine learning model can reasonably be described as a classifier because it classifies “region[s]” of the “the input image 302” as “where the at least one object of interest 302A can be located” and the second machine learning model is a classifier, “The trained DNN2 can perform the classification”; see also Chen, Pg. 3, Para. 4, “we learn stronger classification modules in both RPN and RCN using the knowledge distillation framework”, where “both RPN and RCN” comprise “classification modules”),
wherein the first machine learning model is configured to output a plurality of first classes (Arcadu, Fig. 3 and Arcadu, Para. [0045], “At 304, a first deep neural network DNN1 can be executed. The trained DNN1 can be fed with the input image 302 to generate a region proposal 304A within the image where the at least one object of interest 302A can be located. The trained DNN1 can be able to generate several region proposals within the image, one for each object of interest, if several objects of interest, isolated or overlapping, are present in the input image., where the “several region proposals” are a plurality of first predictions for a plurality of portions of the media content, which are within the broadest reasonable interpretation of first classes because they classify the region as either “a region proposal 304A within the image where the at least one object of interest 302A can be located” or not, which in view of Chen, is an output based on a “classification module[]” of the first machine learning model, see Chen, Pg. 3, Para. 4, “we learn stronger classification modules in both RPN and RCN using the knowledge distillation framework”, where “both RPN and RCN” comprise “classification modules”),
wherein the second machine learning model is configured to output a plurality of second classes, wherein the plurality of first classes is a subset of the plurality of second classes (Arcadu, Fig. 3 and Arcadu, Para. [0045], “The trained DNN1 can be able to generate several region proposals within the image, one for each object of interest, if several objects of interest, isolated or overlapping, are present in the input image . . . The trained DNN2 can perform the classification of the one or more objects of interest in each Region of Interest (RoI), and can return a class 306A for the object of interest”, where, when “several objects of interest . . . are present in the input image”, the second machine learning model, “The trained DNN2”, outputs a plurality of second classes, “return a class 306A for the object of interest”, for respective portions of the plurality of portions, “objects of interest in each Region of Interest (RoI)”, such that the plurality of first classes, “each Region of Interest (RoI)” is a subset of the plurality of second classes, “in each Region of Interest (RoI), and can return a class 306A for the object of interest”).
The reasons for obviousness were discussed in regard to the rejection of Claim 1 and remain applicable here.
Regarding Claim 9, Arcadu teaches a non-transitory computer readable medium having stored thereon instructions for causing a processing circuitry to execute a process, the process comprising: . . . (Fig. 2 and Para. [0037] – [0041], “FIG. 2 depicts a block diagram that illustrates an exemplary data processing apparatus . . . The data processing apparatus 102 can include . . . a processor 204, a memory 206 . . . The processor 204 can be communicatively coupled with the memory 206 . . . the processor 204 can comprise suitable logic, circuitry and interfaces that can be configured to execute programs stored in the memory 206. The programs can correspond to sets of instructions for image processing operations . . . Examples of the processor 204 can include, but are not limited to, Graphical Processing Units (GPUs), a Central Processing Units (CPUs), motherboards, network cards . . . the memory 206 can include, but are not limited to, Random Access Memory (RAM), Read Only Memory (ROM), Hard Disk Drive (HDD), Solid State Drive (SDD) and/or other memory systems”, where a non-transitory computer readable medium, “the memory 206 can include, but are not limited to, Random Access Memory (RAM), Read Only Memory (ROM), Hard Disk Drive (HDD), Solid State Drive (SDD) and/or other memory systems”, stores “instructions” for causing a processing circuitry, “the processor 204 can include, but are not limited to, Graphical Processing Units (GPUs), a Central Processing Units (CPUs), motherboards, network cards”, to execute a process, “the processor 204 . . . can be configured to execute programs . . . The programs can correspond to sets of instructions for image processing operations”).
The remaining limitations are substantially the same as limitations of Claim 1, therefore it is rejected under the same rationale.
Regarding Claim 10, Arcadu teaches a system for visual content processing, comprising: a processing circuitry; and a memory, the memory containing instructions that, when executed by the processing circuitry, configure the system to: . . . (Fig. 2 and Para. [0037] – [0040], “FIG. 2 depicts a block diagram that illustrates an exemplary data processing apparatus . . . The data processing apparatus 102 can include . . . a processor 204, a memory 206 . . . The processor 204 can be communicatively coupled with the memory 206 . . . the processor 204 can comprise suitable logic, circuitry and interfaces that can be configured to execute programs stored in the memory 206. The programs can correspond to sets of instructions for image processing operations . . . Examples of the processor 204 can include, but are not limited to, Graphical Processing Units (GPUs), a Central Processing Units (CPUs), motherboards, network cards”, where an “apparatus” comprising: a processing circuitry, “the processor 204 can include, but are not limited to, Graphical Processing Units (GPUs), a Central Processing Units (CPUs), motherboards, network cards”; and “a memory” containing “instructions” that, when executed by the processing circuitry, perform operations, “the processor 204 . . . can be configured to execute programs . . . The programs can correspond to sets of instructions for image processing operations”, such that the “system” for visual processing, “Systems . . . for detection, segmentation and characterization of isolated or overlapping object instances in digital images”, that “include[s]” the “apparatus” comprises the above discussed components to configure it to perform the above discussed operations, see Abstract, “Systems and methods disclosed herein relate generally to systems and methods for detection, segmentation and characterization of isolated or overlapping object instances in digital images”; Fig. 1; and Para. [0027], “FIG. 1, the system 100 can include a data processing apparatus”).
The remaining limitations are substantially the same as limitations of Claim 1, therefore it is rejected under the same rationale.
Regarding Claim 14, the additional elements of the dependent claim are substantially the same as limitations of Claim 5, therefore it is rejected under the same rationale.
Regarding Claim 15, the additional elements of the dependent claim are substantially the same as limitations of Claim 6, therefore it is rejected under the same rationale.
Regarding Claim 17, the additional elements of the dependent claim are substantially the same as limitations of Claim 8, therefore it is rejected under the same rationale.
Regarding Claim 18, Arcadu teaches a method for visual content processing, comprising (Abstract, “methods disclosed herein relate generally to systems and methods for detection, segmentation and characterization of isolated or overlapping object instances in digital images”, where the “methods” are for visual content processing, “for detection, segmentation and characterization . . . in digital images”; see also Fig. 3):
applying a first machine learning model to media content . . . (Abstract, “Systems and methods disclosed herein relate generally to systems and methods for detection, segmentation and characterization of isolated or overlapping object instances in digital images”; Fig. 3 and Para. [0045], “The trained DNN1 can be fed with the input image 302”, where a first machine learning model, “DNN1” is applied media content, “digital images” as “input”)
selecting a subset of the media content based on the plurality of predictions output by the first machine learning model (Fig. 3 and Para. [0045], “The trained DNN1 can be fed with the input image 302 to generate a region proposal 304A within the image where the at least one object of interest 302A can be located. The trained DNN1 can be able to generate several region proposals within the image, one for each object of interest, if several objects of interest, isolated or overlapping, are present in the input image”, where a subset of the media content, “a region proposal 304A within the image”, is selected based on the plurality of predictions output by the first machine learning model, “The trained DNN1 can be able to generate several region proposals within the image, one for each object of interest, if several objects of interest . . . , are present in the input image”); and
providing the selected subset of media content as inputs to a second machine learning model . . . (Fig. 3 and Para. [0045], “At 306, a second deep neural network DNN2 can be executed on the at least one Region of Interest (RoI) 304A defined at 304”, where a second machine learning model, “a second deep neural network DNN2”, is provided the selected subset of media content as inputs, “executed on the at least one Region of Interest (RoI) 304A defined at 304”).
The remaining limitations are substantially the same as limitations of Claim 1, therefore it is rejected under the same rationale.
Regarding Claim 19, Arcadu in view of Chen teach the method of claim 18, wherein the media content includes a plurality of images, further comprising (Arcadu, Fig. 3 and Arcadu, Abstract, “methods disclosed herein relate generally to systems and methods for detection, segmentation and characterization of isolated or overlapping object instances in digital images”, where the media content includes a plurality of images, which analyzed on image-by-image basis as “input” to the method depicted in figure 3):
cropping from among the plurality of images based on the plurality of predictions output by the first machine learning model in order to create a plurality of cropped images (Arcadu, Fig. 3 and Arcadu, Para. [0045], “The trained DNN1 can be fed with the input image 302 to generate a region proposal 304A within the image where the at least one object of interest 302A can be located . . . At 306, a second deep neural network DNN2 can be executed on the at least one Region of Interest (RoI) 304A defined at 304”, where a “a region proposal 304A within the image” is within the broadest reasonable interpretation of a cropping to create a plurality of cropped images because it removes the data outside of the “at least one Region” from downstream processes, “a second deep neural network DNN2 can be executed on the at least one Region of Interest (RoI) 304A defined at 304”, which is based on the plurality of predictions output by the first machine learning model, “The trained DNN1 can . . . generate a region proposal 304A”; see also Arcadu, Fig. 3 and Arcadu, Abstract, “methods disclosed herein relate generally to systems and methods for detection, segmentation and characterization of isolated or overlapping object instances in digital images”, where the media content includes a plurality of images, which analyzed on image-by-image basis as “input” to the method depicted in figure 3),
wherein the selected subset of the media content includes the plurality of cropped images (Arcadu, Fig. 3 and Arcadu, Para. [0045], “At 306, a second deep neural network DNN2 can be executed on the at least one Region of Interest (RoI) 304A defined at 304”, where the second machine learning model, “a second deep neural network DNN2”, is provided the selected subset of media content as inputs, “executed on the at least one Region of Interest (RoI) 304A defined at 304”, which includes the plurality of cropped images, “can be executed on the at least one Region of Interest (RoI) 304A defined at 304”; see also Arcadu, Fig. 3 and Arcadu, Abstract, “methods disclosed herein relate generally to systems and methods for detection, segmentation and characterization of isolated or overlapping object instances in digital images”, where the media content includes a plurality of images, which analyzed on image-by-image basis as “input” to the method depicted in figure 3).
Claims 2 and 11 are rejected under 35 U.S.C. 103 as being unpatentable over Arcadu in view of Chen and Li et al. (hereinafter Li) (Pat. App. Pub. No. US 2021/0407090 A1).
Regarding Claim 2, Arcadu in view of Chen teach the method of claim 1, further comprising: training the student model using the teacher model (Chen, Pg. 4, Para. 2, “Let t be the teacher model . . . The student s is trained to optimize the following loss function . . . It is known that a deep teacher can better fit to the training data and perform better in test scenarios. The soft labels contain information about the relationship between different classes as discovered by teacher. By learning from soft labels, the student network inherits such hidden information”, wherein the student model is trained, “The student s is trained to optimize the following loss function”, using the teacher model, “the relationship between different classes as discovered by teacher. By learning from soft labels, the student network inherits such hidden information”),
wherein a domain used by the student model is a subset of a domain used by the teacher model (Chen, Pg. 2, Para. 2, “we focus on transferring knowledge within the same domain (images of the same dataset) with no additional data or labels, as opposed other works that might rely on data from other domains (such as high-quality and low-quality image domains, or image and depth domains)”, wherein the domain used by the student model is the same domain used by the teacher model, “we focus on transferring knowledge within the same domain (images of the same dataset)”, such that the student domain is a subset of the teacher domain because all elements of the student domain are in the teacher domain; see also Chen, Pg. 2, Para. 3, “, a teacher bounded regression loss for knowledge distillation (Section 3.3) and adaptation layers for hint learning that allows the student to better learn from the distribution of neurons in intermediate layers of the teacher”);
. . . as the first machine learning model . . . (Chen, Pg. 3, Para. 4, “We learn strong but efficient student object detectors by using the knowledge of a high capacity teacher detection network for all the three components”; see also Chen, Pg. 4, Para. 2, “Let t be the teacher model . . . The student s is trained to optimize the following loss function . . . It is known that a deep teacher can better fit to the training data and perform better in test scenarios. The soft labels contain information about the relationship between different classes as discovered by teacher. By learning from soft labels, the student network inherits such hidden information”).
The reasons for obviousness were discussed in regard to the rejection of Claim 1 and remain applicable here.
Arcadu in view of Chen do not explicitly disclose . . . and sending, from a second system to a first system, the trained student model for deployment . . . at the first system, wherein the first system is remote from the second system.
However, Li teaches . . . [training the student model using the teacher model] and sending, from a second system to a first system, the trained student model for deployment . . . at the first system, wherein the first system is remote from the second system (Abstract, “Training the student model includes using selected outputs of the specialized teacher model. The method further includes deploying the trained student model to perform visual object instance segmentation in an external device”; Fig. 2; and Para. [0051], “In some embodiments, the training operations 206 and 208 can be performed in a cloud-based environment, such as when the training operations 206 and 208 are performed using one or more computing devices (such as one or more servers 106) in a computing cloud”, where the student model is trained using the teacher model, “Training the student model includes using selected outputs of the specialized teacher model”, and sent as a trained model for deployment, “deploying the trained student model”, from a second system, “a cloud-based environment”, to a first system, “an external device”, wherein the first system is remote from the second system, “a cloud-based environment” being remote from an “external device”).
Before the effective filing date of the invention, it would have been obvious to one of ordinary skill in the art to combine the training the student model as the first machine learning model using the teacher model, wherein a domain used by the student model is a subset of a domain used by the teacher model of Arcadu in view of Chen with the training the student model using the teacher model and sending, from a second system to a first system, the trained student model for deployment at the first system, wherein the first system is remote from the second system of Li in order to utilize systems with abundant resources to generate the student model, which can then be deployed to a resource constrained device to perform tasks with increased accuracy (Li, Para. [0030], “due to the complexity of the visual object instance segmentation task, high-accuracy visual object instance segmentation algorithms often rely on large and complicated deep neural networks. As a result, these approaches typically cannot be used on mobile or edge devices, which are often constrained both in terms of processing power and memory . . . which prevents visual object instance segmentation from being performed accurately on those devices” and Li, Para. [0094], “The online platform 1106 may represent a server, a cloud-based hosting environment, or other device or system in which memory and processing resources are more abundant (at least relative to the external devices 212 described above)”).
Regarding Claim 11, the additional elements of the dependent claim are substantially the same as limitations of Claim 2, therefore it is rejected under the same rationale.
Claims 3-4 and 12-13 are rejected under 35 U.S.C. 103 as being unpatentable over Arcadu in view of Chen, Li, and Siemionow et al. (hereinafter Siemionow) (Pat. App. Pub. No. US 2022/0335687 A1).
Regarding Claim 3, Arcadu in view of Chen and Li teach the method of claim 2, wherein the media content is second media content (Chen, Pg. 6, Para. 1, “we use the training/validation split introduced by [39] and [24] for analysis”, where the “training” data used for student model training, Chen, Pg. 4, Para. 2, “The student s is trained to optimize the following loss function”, is separate from the “validation”, such that the media content “input” data, can reasonable be described as second media content, see Arcadu, Fig. 3 and Arcadu, Para. [0045], “The trained DNN1 can be fed with the input image 302 to generate a region proposal 304A within the image where the at least one object of interest 302A can be located . . . At 306, a second deep neural network DNN2 can be executed on the at least one Region of Interest (RoI) 304A defined at 304”,),
wherein training the student model using the teacher model further comprises (Chen, Pg. 4, Para. 2, “Let t be the teacher model . . . The student s is trained to optimize the following loss function . . . It is known that a deep teacher can better fit to the training data and perform better in test scenarios. The soft labels contain information about the relationship between different classes as discovered by teacher. By learning from soft labels, the student network inherits such hidden information”, wherein the student model is trained, “The student s is trained to optimize the following loss function”, using the teacher model, “the relationship between different classes as discovered by teacher. By learning from soft labels, the student network inherits such hidden information”):
applying the student model to first media content in order to output a plurality of student predictions . . . the plurality of student predictions . . . applying the teacher model to the selected subset of the first media content in order to output a plurality of teacher predictions (Chen, Pg. 4, Para. 2, “Conventional use of knowledge distillation has been proposed for training classification networks, where predictions of a teacher network are used to guide the training of a student model. Suppose we have dataset {xi,yi}, i = 1,2,...,n where xi ∈ I is the input image and yi ∈ Y is its class label. Let t be the teacher model, with Pt = [SOFTMAX] its prediction and Zt the final score output. Here, T is a temperature parameter (normally set to 1). Similarly, one can define Ps = [SOFTMAX] for the student network s. The student s is trained to optimize the following loss function” and Chen, Pg. 3, Para. 3, “returns the detection score as well as a spatial adjustment vector for each object proposal”, where both the “student” model and the “teacher” model are applied to a selected subset of the first media content, “dataset {xi,yi}, i = 1,2,...,n where xi ∈ I is the input image”, in order to output a plurality of student predictions, “Ps” for “each object proposal”, and a plurality of teacher predictions, “Pt” “predictions of a teacher”); and
tuning the student model based on the plurality of teacher predictions (Chen, Pg. 4, Para. 2, “Let t be the teacher model . . . The student s is trained to optimize the following loss function . . . It is known that a deep teacher can better fit to the training data and perform better in test scenarios. The soft labels contain information about the relationship between different classes as discovered by teacher. By learning from soft labels, the student network inherits such hidden information”, wherein the student model is tuned, “The student s is trained to optimize the following loss function”, using the plurality of “teacher” predictions, “soft labels”,; see also Chen, Pg. 3, Para. 4, “We learn strong but efficient student object detectors by using the knowledge of a high capacity teacher detection network for all the three components”),
wherein the first machine learning model is created based on the tuned student model (Chen, Pg. 3, Para. 4, “We learn strong but efficient student object detectors by using the knowledge of a high capacity teacher detection network for all the three components”; see also Chen, Pg. 4, Para. 2, “Let t be the teacher model . . . The student s is trained to optimize the following loss function . . . It is known that a deep teacher can better fit to the training data and perform better in test scenarios. The soft labels contain information about the relationship between different classes as discovered by teacher. By learning from soft labels, the student network inherits such hidden information”).
The reasons for obviousness were discussed in regard to the rejection of Claim 1 and the rejection of Claim 2 and remain applicable here.
Arcadu in view of Chen and Li do not explicitly disclose . . . selecting a subset of the first media content based on . . . .
However, Siemionow teaches . . . selecting a subset of the first media content based on [predictions from the model in training] . . . (Para. [0051], “The training process may include periodic check of the prediction accuracy using a held out input data set (the validation set) not included in the training data. If the check reveals that the accuracy on the validation set is better than the one achieved during the previous check, the complete neural network weights are stored for further use. The early stopping function may terminate the training if there is no improvement observed during the last CH checks”, where a subset of the first media content, “the training data” is selected based on predictions form the model in training, “The early stopping function may terminate the training if there is no improvement observed during the last CH checks”).
Before the effective filing date of the invention, it would have been obvious to one of ordinary skill in the art to combine the applying of the student model to first media content in order to output a plurality of student predictions and applying the teacher model to a selected subset of the first media content in order to output a plurality of teacher predictions of Arcadu in view of Chen and Li with the selecting a subset of the first media content based on predictions from the model in training of Siemionow in order to improve student model training through increased training accuracy and reduction of ineffective computations for the student training (Chen, Pg. 7, Para. 2, “Designing new models is one option. But it often requires significant labor towards design and training” with Siemionow, Para. [0045] – [0051], “During training, the following means of improving the training accuracy can be used: . . . early stopping . . . The training process may include periodic check of the prediction accuracy using a held out input data set (the validation set) not included in the training data. If the check reveals that the accuracy on the validation set is better than the one achieved during the previous check, the complete neural network weights are stored for further use. The early stopping function may terminate the training if there is no improvement observed during the last CH checks”).
Regarding Claim 4, Arcadu in view of Chen, Li, and Siemionow teach the method of claim 3, wherein tuning the student model using the plurality of teacher predictions further comprises: generating a plurality of teacher prediction labels based on the plurality of teacher predictions, wherein the plurality of teacher prediction labels is used to tune the student model (Chen, Pg. 4, Para. 2, “Let t be the teacher model . . . The student s is trained to optimize the following loss function . . . It is known that a deep teacher can better fit to the training data and perform better in test scenarios. The soft labels contain information about the relationship between different classes as discovered by teacher. By learning from soft labels, the student network inherits such hidden information”, wherein the student model is tuned, “The student s is trained to optimize the following loss function”, using the plurality of “teacher” predictions, “soft labels”, which comprises generating a plurality of teacher prediction labels based on the plurality of teacher predictions, “The soft labels contain information about the relationship between different classes as discovered by teacher”, wherein the plurality of teacher prediction labels, “soft labels”, is used to tune the student model, “The student s is trained to optimize the following loss function”; see also Chen, Pg. 3, Para. 4, “We learn strong but efficient student object detectors by using the knowledge of a high capacity teacher detection network for all the three components”).
The reasons for obviousness were discussed in regard to the rejection of Claim 1 and remain applicable here.
Regarding Claim 12, the additional elements of the dependent claim are substantially the same as limitations of Claim 3, therefore it is rejected under the same rationale.
Regarding Claim 13, the additional elements of the dependent claim are substantially the same as limitations of Claim 4, therefore it is rejected under the same rationale.
Claims 7 and 16 are rejected under 35 U.S.C. 103 as being unpatentable over Arcadu in view of Chen and Gazzetti et al. (hereinafter Gazzetti) (Pat. App. Pub. No. US 2021/0065063 A1).
Regarding Claim 7, Arcadu in view of Chen teach the method of claim 1, wherein the first machine learning model is applied . . . with respect to a source of the media content, wherein the second machine learning model is deployed . . . (Arcadu, Fig. 3 and Arcadu, Para. [0045], “The trained DNN1 can be fed with the input image 302 to generate a region proposal 304A within the image where the at least one object of interest 302A can be located . . . At 306, a second deep neural network DNN2 can be executed on the at least one Region of Interest (RoI) 304A defined at 304”, where the first machine learning model is applied with respect to a source of the media content, “The trained DNN1 can be fed with the input image 302”, wherein the second machine learning model is deployed, “At 306, a second deep neural network DNN2 can be executed”).
Arcadu in view of Chen do not explicitly disclose . . . by an edge device deployed locally . . . remotely from the edge device.
However, Gazzetti teaches . . . [first machine learning model applied] by an edge device deployed locally . . . [and a second machine learning model deployed] remotely from the edge device (Fig. 16 and Para. [0105], “The edge service 1604 may be implemented in any suitable computing device (e.g., perhaps integrated into the same device as the rest API 1602) and is configured process the input data (e.g., the frame buffer 1608) with a (first) machine learning model (e.g., an object classification model). If determined to be appropriate, the edge service 1604 may also interact with the cloud platform (or remote cloud platform) 1606 to request that the input data be processed with another (e.g. a second) machine learning model”).
Before the effective filing date of the invention, it would have been obvious to one of ordinary skill in the art to combine the first machine learning model that is applied with respect to a source of the media content and the second machine learning model is deployed of Arcadu in view of Chen with the first machine learning model applied by an edge device deployed locally and a second machine learning model deployed remotely from the edge device in order to utilize the increased resources of the cloud service to perform the more computationally intensive operations of the second machine learning model with increased accuracy (compare Gazzetti, Para. [0105], “As such, the cloud platform 1606 may include, for example, remote services/systems deployed in a cloud infrastructure, which are configured to process the input data with the second (or third, fourth, etc.) machine learning model. It should be noted that because of the potentially (nearly) unlimited resource constraints of such a service (e.g., on the cloud), the accuracy of inferences generated by such a system may exceed that of the edge service 1604 (e.g., implemented locally, on a single device, etc.)” with Chen, Pg. 3, Para. 3, “we adopt the Faster-RCNN [32] as the object detection framework. Faster-RCNN is composed of three modules: 1) A shared feature extraction through convolutional layers, 2) a region proposal network (RPN) that generates object proposals, and 3) a classification and regression network (RCN) that returns the detection score as well as a spatial adjustment vector for each object proposal. Both the RCN and RPN use the output of 1) as features, RCN also takes the result of RPN as input” and Arcadu, Para. [0045], “In several embodiments, the trained DNN2 can comprise a chain of several trained DNNs to perform tasks sequentially or in parallel”; see also Arcadu, Para. [0018], “Convolutional Neural Networks (CNNs) are particularly suited for image recognition tasks. CNNs are built such that the processing units in the early layers learn to activate in response to simple local features, for example patterns at particular orientations or edges, while units in the deeper layers combine the low-level features into more complex patterns”).
Regarding Claim 16, the additional elements of the dependent claim are substantially the same as limitations of Claim 7, therefore it is rejected under the same rationale.
Response to Arguments
Applicant's arguments filed on June 17th, 2026, have been fully considered. Each argument is addressed in detail below.
I. Applicant argues the rejections of claims 1-19, under 35 USC § 101, should be withdrawn (Applicant’s Remarks, 06/17/2026, Pg. 9-13, Section “Claim Rejections – 35 U.S.C. § 101”).
In response to Applicant’s amendments, the rejections of claims 1-17, under 35 USC § 101, have been withdrawn. As a result, Applicant’s arguments, which are directed toward the subject matter eligibility of Claim 1 as a representative example, are rendered moot with regard to claims 1-17.
The rejections of claims 18-19, under 35 USC § 101, are maintained. Notably, neither claim 18 nor claim 19 positively recite elements related to domain subsets (compare Claim 1, “wherein a domain used by the first machine learning model is a subset of a domain used by the second machine learning model” with Claim 18, “wherein a domain used by the first machine learning model is smaller than a domain used by the second machine learning model”), which applicant argues provides for a technical improvement that integrates the judicial exception into a practical application (Applicant’s Remarks, Pg. 9-13). As a result, Applicant’s arguments are rendered moot.
However, in the interests of compact prosecution and clarity, it is worth elaborating on the analysis related to the subject matter eligibility of independent claim 18 and its dependent claims. Specifically, the claims positively recites elements, such as the use of varied-domain machine learning model in sequence, that have broad applicability across many fields of machine learning endeavor (see MPEP 2106.05(f), “A claim having broad applicability across many fields of endeavor may not provide meaningful limitations that integrate a judicial exception into a practical application or amount to significantly more”), such as to fail to reflect the alleged improvements (see MPEP 2106.04(d)(1)), “the claim must be evaluated to ensure that the claim itself reflects the disclosed improvement”). As a result, the additional elements do not integrate the abstract ideas into a practical application or amount to significantly more.
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure.
Sharma et al. (Pat. App. Pub. No. US 2020/0050846 A1) discloses additional technical details relating to training a student model using a teacher model.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to MATTHEW BRYCE GOLAN whose telephone number is (571)272-5159. The examiner can normally be reached Monday through Friday, 8:00 AM to 5:00 PM ET.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Alexey Shmatov can be reached at (571) 270-3428. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/MATTHEW BRYCE GOLAN/Examiner, Art Unit 2123
/ALEXEY SHMATOV/Supervisory Patent Examiner, Art Unit 2123