DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Election/Restrictions
Claims 2, 3, and 19 have been withdrawn from further consideration pursuant to 37 CFR 1.142(b) as being drawn to a nonelected Specie, there being no allowable generic or linking claim. Election was made without traverse in the reply filed on March 20, 2026.
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claim(s) 1, 5-6, 8-17 is/are rejected under 35 U.S.C. 103 as being unpatentable over Avinash et al. (WO 2025/101175 A1) in view of Hannuksela (WO 2025/008694 A1).
Regarding claim 1 Avinash discloses one or more processors comprising:
one or more circuits (one or more processors – [18]) to:
receive a plurality of frames (plurality of images – [4]);
determine that at least one frame of the plurality of frames is to be provided as input to a machine-learning model; generate an indication that the at least one frame is to be provided as input to the machine-learning model (the computing system can select a small number of borderline or difficult-to-label image examples and ask the user to label only the selected images – [47]; the labelled image dataset can be used to train a machine-learned image classifier – [48]).
However, fails to explicitly disclose receiving a plurality of frames from a capture device capturing a video stream and to generate an encoded bitstream for the video stream, the encoded bitstream including encoded data for the plurality of frames and the indication.
In his disclosure Hannuksela teaches receiving a plurality of frames from a capture device capturing a video stream (camera – [0069, 0073]) and to generate an encoded bitstream for the video stream, the encoded bitstream including encoded data for the plurality of frames and the indication (an encoder that encodes one or more indication that indicate images that are used as input for a neural network – [0424]).
It would have been obvious to a person with ordinary skill in the art, before the effective filing date of the claimed invention, to incorporate the teachings of Hannuksela into the teachings of Avinash because such incorporation yields the predictable result of reducing the processing time.
Regarding claim 5 Avinash discloses the one or more processors of claim 1. However, fails to explicitly disclose wherein the one or more circuits are to: generate the indication to include a binary value indicating that the at least one frame is to be provided as input to the machine-learning model.
In his disclosure Hannuksela teaches generate the indication to include a binary value indicating that the at least one frame is to be provided as input to the machine-learning model (indications having value of “1” – [0352-0353]).
It would have been obvious to a person with ordinary skill in the art, before the effective filing date of the claimed invention, to incorporate the teachings of Hannuksela into the teachings of Avinash because such incorporation yields the predictable result of reducing the processing time.
Regarding claim 6 Avinash discloses the one or more processors of claim 1. However, fails to explicitly disclose wherein the one or more circuits are to: generate the indication to include supplemental enhancement information (SEI) indicating that the at least one frame is to be provided as input to the machine-learning model.
In his disclosure Hannuksela teaches generate the indication to include supplemental enhancement information (SEI) indicating that the at least one frame is to be provided as input to the machine-learning model (indication(s) identifying the input pictures for an NNPF in an NNPF group are encoded into and/or decoded from a postfilter group SEI message or alike – [0416]).
It would have been obvious to a person with ordinary skill in the art, before the effective filing date of the claimed invention, to incorporate the teachings of Hannuksela into the teachings of Avinash because such incorporation yields the predictable result of reducing the processing time.
Regarding claim 8 Avinash discloses the one or more processors of claim 1. However, fails to explicitly disclose wherein the one or more circuits are to: transmit the encoded bitstream to a receiver system, causing the receiver system to decode the encoded bitstream and provide the at least one frame as input to the machine-learning model.
In his disclosure Hannuksela teaches transmit the encoded bitstream to a receiver system, causing the receiver system to decode the encoded bitstream and provide the at least one frame as input to the machine-learning model (receiving apparatus in Figure 13).
It would have been obvious to a person with ordinary skill in the art, before the effective filing date of the claimed invention, to incorporate the teachings of Hannuksela into the teachings of Avinash because such incorporation yields the predictable result of reducing the processing time.
Regarding claim 9 Avinash discloses the one or more processors of claim 8. However, fails to explicitly disclose wherein the one or more circuits are to: transmit the encoded bitstream according to a real time streaming protocol (RTSP).
In his disclosure Hannuksela discloses transmit the encoded bitstream according to a real time streaming protocol (RTSP) (RTSP – [0101]).
It would have been obvious to a person with ordinary skill in the art, before the effective filing date of the claimed invention, to incorporate the teachings of Hannuksela into the teachings of Avinash because such incorporation yields the predictable result of reducing the processing time.
Regarding claim 10 Avinash discloses the one or more processors of claim 1, wherein the one or more processors are comprised in at least one of:
a control system for an autonomous or semi-autonomous machine; a perception system for an autonomous or semi-autonomous machine; a system for performing simulation operations; a system for performing digital twin operations; a system for performing light transport simulation; a system for performing collaborative content creation for 3D assets; a system for performing deep learning operations; a system implemented using an edge device; a system implemented using a robot; a system for performing conversational AI operations; a system for performing generative AI operations using a large language model (LLM); a system for performing generative AI operations using a video language model (VLM); a system for generating synthetic data; a system incorporating one or more virtual machines (VMs); a system implemented at least partially in a data center; or a system implemented at least partially using cloud computing resources (vision language models (VLMs) – [43]).
Regarding claim 11 Avinash discloses a system, comprising: one or more processors to: receive an encoded bitstream of a video stream; decode the encoded bitstream to obtain a plurality of frames and an indication that at least one frame of the plurality of frames is to be provided as input to a machine- learning model; and provide the at least one frame as input to the machine-learning model according to the indication (refer to rejection of claim 1. It is noted Hannuksela teaches the decoding of an encoded bitstream – Figure 13).
Regarding claim 12 Avinash discloses the system of claim 11. However, fails to explicitly disclose wherein the one or more processors are to: retrieve the encoded bitstream of the video stream from a database.
In his disclosure Hannuksela teaches retrieve the encoded bitstream of the video stream from a database (Hannuksela teaches it is known in the art for storing bitstreams in servers – [0086, 0159, 0160]).
It would have been obvious to a person with ordinary skill in the art, before the effective filing date of the claimed invention, to incorporate the teachings of Hannuksela into the teachings of Avinash because such incorporation yields the predictable result of reducing the processing time.
Regarding claim 13 Avinash discloses the system of claim 11. However, fails to explicitly disclose wherein the one or more processors are to: generate metadata by decoding the encoded bitstream, the metadata comprising the indication that the at least one frame is to be provided as input to the machine-learning model.
In his disclosure Hannuksela discloses generate metadata by decoding the encoded bitstream, the metadata comprising the indication that the at least one frame is to be provided as input to the machine-learning model (decoding of an encoded bitstream – Figure 13; an encoder that encodes one or more indication that indicate images that are used as input for a neural network – [0424]).
It would have been obvious to a person with ordinary skill in the art, before the effective filing date of the claimed invention, to incorporate the teachings of Hannuksela into the teachings of Avinash because such incorporation yields the predictable result of reducing the processing time.
Regarding claim 14 Avinash discloses the system of claim 11, wherein the one or more processors are to: update the machine-learning model using the at least one frame (a training system 112 can provide model updates 460 based on the newly labeled images – [98]).
Regarding claim 15 Avinash discloses the system of claim 11, wherein the machine-learning model comprises a video language model (VLM) (vison language models – [43]).
Regarding claim 16 Avinash discloses the system of claim 11, wherein the system is comprised in at least one of: a control system for an autonomous or semi-autonomous machine; a perception system for an autonomous or semi-autonomous machine; a system for performing simulation operations; a system for performing digital twin operations; a system for performing light transport simulation; a system for performing collaborative content creation for 3D assets; a system for performing deep learning operations; a system implemented using an edge device; a system implemented using a robot; a system for performing conversational AI operations; a system for performing generative AI operations using a large language model (LLM);a system for performing generative AI operations using a video language model (VLM);a system for generating synthetic data; a system incorporating one or more virtual machines (VMs);a system implemented at least partially in a data center; or a system implemented at least partially using cloud computing resources (vision language models (VLMs) – [43]).
Claim 17 corresponds to the method performed by the processor of claim 1. Therefore, claim 17 is being rejected on the same basis as claim 1.
Claim(s) 4, 7, 18, and 20 is/are rejected under 35 U.S.C. 103 as being unpatentable over Avinash et al. (WO 2025/101175 A1) in view of Hannuksela (WO 2025/008694 A1) further in view of Gao et al. (US 2026/0019579).
Regarding claim 4 Avinash discloses the one or more processors of claim 1. However, fails to explicitly disclose wherein the one or more circuits are to: determine, using a second machine-learning model, that the at least one frame depicts an object of interest; and determine that the at least one frame is to be provided as input to the machine-learning model responsive to determining that the at least one frame depicts the object of interest.
In his disclosure Gao teaches determine, using a second machine-learning model, that the at least one frame depicts an object of interest; and determine that the at least one frame is to be provided as input to the machine-learning model responsive to determining that the at least one frame depicts the object of interest (Faster-RCNN is known as a model in which a region based convolutional neural network (R-CNN), which is a region-based object detection model is sped up. In Faster-RCNN, a plurality of feature maps having different sizes in each hierarchical layer are generated by performing convolution processing on an input image of a processing target using the first neural network (feature pyramid network) of a plurality of hierarchical layers. Then, by applying the generated feature map with an RP model using the second neural network (region proposal network), an ROI region is extracted from the feature map, and image recognition is performed on the extracted ROI region – [0039]).
It would have been obvious to a person with ordinary skill in the art, before the effective filing date of the claimed invention, to incorporate the teachings of Gao into the teachings of Avinash because such incorporation improves the compression efficiency.
Regarding claim 7 Avinash discloses the one or more processors of claim 6. However, fails to explicitly disclose wherein the SEI information includes an indication of at least one object detected in the frame.
In his disclosure Gao teaches the SEI information includes an indication of at least one object detected in the frame (syntax information or index information designating a generation method of the unit feature map and a generation method of the picture may be decoded from a supplemental enhancement information (SEI) region or another header region of the bitstream – [0116]).
It would have been obvious to a person with ordinary skill in the art, before the effective filing date of the claimed invention, to incorporate the teachings of Gao into the teachings of Avinash because such incorporation improves the compression efficiency.
Regarding claim 18 Avinash discloses the method of claim 17. However, fails to explicitly disclose wherein the at least one attribute includes at least one of an object detected in the at least one frame or a temporal activity detected in the at least one frame.
In his disclosure Gao teaches the at least one attribute includes at least one of an object detected in the at least one frame or a temporal activity detected in the at least one frame (an ROI region is extracted from the feature map, and image recognition is performed on the extracted ROI region – [0039]).
It would have been obvious to a person with ordinary skill in the art, before the effective filing date of the claimed invention, to incorporate the teachings of Gao into the teachings of Avinash because such incorporation improves the compression efficiency.
Claim 20 corresponds to the method performed by the apparatus of claim 4. Therefore, claim 20 is being rejected on the same basis as claim 4.
Conclusion
Any inquiry concerning this communication or earlier communications from the examiner should be directed to MARIA E VAZQUEZ COLON whose telephone number is (571)270-1103. The examiner can normally be reached M-F 7:30 AM-3:30 PM.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, CHRISTOPHER S KELLEY can be reached at (571)272-7331. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/MARIA E VAZQUEZ COLON/ Examiner, Art Unit 2482