Prosecution Insights
Last updated: August 17, 2026
Application No. 18/879,761

A METHOD, AN APPARATUS AND A COMPUTER PROGRAM PRODUCT FOR VIDEO CODING

Non-Final OA §102§112
Filed
Dec 28, 2024
Priority
Jun 30, 2022 — FI 20225599 +1 more
Examiner
DANG, PHILIP
Art Unit
2488
Tech Center
2400 — Computer Networks
Assignee
Nokia Corporation
OA Round
1 (Non-Final)
78%
Grant Probability
Favorable
1-2
OA Rounds
1y 0m
Est. Remaining
99%
With Interview

Examiner Intelligence

Grants 78% — above average
78%
Career Allowance Rate
385 granted / 496 resolved
+19.6% vs TC avg
Strong +30% interview lift
Without
With
+30.4%
Interview Lift
resolved cases with interview
Typical timeline
2y 7m
Avg Prosecution
30 currently pending
Career history
535
Total Applications
across all art units

Statute-Specific Performance

§101
5.2%
-34.8% vs TC avg
§103
53.5%
+13.5% vs TC avg
§102
12.5%
-27.5% vs TC avg
§112
25.5%
-14.5% vs TC avg
Black line = Tech Center average estimate • Based on career data from 496 resolved cases

Office Action

§102 §112
DETAILED ACTION Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Information Disclosure Statement The information disclosure statement (IDS), submitted on 5/21/2025, is being considered by the examiner. Claim Interpretation The following is a quotation of 35 U.S.C. 112(f): (f) ELEMENT IN CLAIM FOR A COMBINATION.—An element in a claim for a combination may be expressed as a means or step for performing a specified function without the recital of structure, material, or acts in support thereof, and such claim shall be construed to cover the corresponding structure, material, or acts described in the specification and equivalents thereof. The following is a quotation of pre-AIA 35 U.S.C. 112, sixth paragraph: An element in a claim for a combination may be expressed as a means or step for performing a specified function without the recital of structure, material, or acts in support thereof, and such claim shall be construed to cover the corresponding structure, material, or acts described in the specification and equivalents thereof. The claims in this application are given their broadest reasonable interpretation using the plain meaning of the claim language in light of the specification as it would be understood by one of ordinary skill in the art. The broadest reasonable interpretation of a claim element (also commonly referred to as a claim limitation) is limited by the description in the specification when 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is invoked. As explained in MPEP § 2181, subsection I, claim limitations that meet the following three-prong test will be interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph: (A) the claim limitation uses the term “means” or “step” or a term used as a substitute for “means” that is a generic placeholder (also called a nonce term or a non-structural term having no specific structural meaning) for performing the claimed function; (B) the term “means” or “step” or the generic placeholder is modified by functional language, typically, but not always linked by the transition word “for” (e.g., “means for”) or another linking word or phrase, such as "configured to" or "so that"; and (C) the term “means” or “step” or the generic placeholder is not modified by sufficient structure, material, or acts for performing the claimed function. Absence of the word “means” (or “step”) in a claim creates a rebuttable presumption that the claim limitation is not to be treated in accordance with 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph. The presumption that the claim limitation is not interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is rebutted when the claim limitation recites function without reciting sufficient structure, material or acts to entirely perform the recited function. Claim limitations in this application that use the word “means” (or “step”) are being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, except as otherwise indicated in an Office action. Conversely, claim limitations in this application that do not use the word “means” (or “step”) are not being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, except as otherwise indicated in an Office action. This application includes one or more claim limitations that do not use the word “means,” but are nonetheless being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, because the claim limitations use a generic placeholder that is coupled with functional language without reciting sufficient structure to perform the recited function and the generic placeholder is not preceded by a structural modifier. Such claim limitations are target data units, auxiliary data units in claim 1, 23, and 24; motion-compensated data units in claim 16 and 26. Because these claim limitations are being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, they are being interpreted to cover the corresponding structure described in the specification as performing the claimed function, and equivalents thereof. If applicant does not intend to have these limitations interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, applicant may: (1) amend the claim limitation(s) to avoid it/them being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph (e.g., by reciting sufficient structure to perform the claimed function); or (2) present a sufficient showing that the claim limitations recite sufficient structure to perform the claimed function so as to avoid it/them being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph. Claim Rejections - 35 USC § 102 The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action: A person shall be entitled to a patent unless – (a)(1) the claimed invention was patented, described in a printed publication, or in public use, on sale or otherwise available to the public before the effective filing date of the claimed invention. Claims 16-30 are rejected under 35 U.S.C. 102(a)(1) as being anticipated by Chen (US Patent 11,531,850 B2), (“Chen”). Regarding claim 16, Chen meets the claim limitations, as follows: An apparatus (an apparatus) [Chen: Abstract; col. 112, line 25] comprising: at least one processor (a processor 552, which may be a microprocessor, a multi-core processor, a multithreaded processor, an ultra-low voltage processor, an embedded processor, or other known processing element.) [Chen: col. 14, line 19-22; Figs. 5-7, 64]; and at least one memory (a system memory) [Chen: col. 14, line 41-42; Figs. 5-7, 64] storing instructions ((computer instructions) [Chen: col. 112, line 47]; (Program code, such as code 730 illustrated in FIG. 7, may 30 be applied to input instructions to perform the functions described herein and generate output information.) [Chen: col. 20, line 29-31; Fig. 7]) that, when executed by the at least one processor, cause the apparatus at least to perform ((Processor 600 can execute any type of instructions associated with algorithms, processes, or operations detailed herein. Generally, processor 600 can transform an element or an article (e.g., data) from one state or thing to another state or thing) [Chen: col. 18, line 41-45]; (In another example, some activities outlined herein may be implemented with fixed logic or programmable logic (for example, software and/or computer instructions executed by a processor) and the elements identified herein could be some type of a programmable processor, programmable digital logic (for example, a field programmable gate array (FPGA), an erasable programmable read only memory (EPROM), an electrically erasable programmable read only memory (EEPROM)), an ASIC that includes digital logic, software, code, electronic instructions, flash memory, optical disks, CD-ROMs, DVD ROMs, magnetic or optical cards, other types of machine-readable mediums suitable for storing electronic instructions, or any suitable combination thereof) [Chen: col. 112, line 45-58]): receiving one or more target data units ((In the illustrated example, the original image is first received 2202) [Chen: col. 37, line 11-12; Figs. 13, 19, 22]; (FIG. 17 illustrates an example embodiment of a runtime processing pipeline 1700 for a visual fog architecture. In the illustrated embodiment, for example, a raw stream of visual data 1701 (e.g., video or images) captured by cameras or visual sensors in a visual fog architecture is provided as input to a stream ingress framework 1702) [Chen: col. 31, line 33-38; Fig. 17]) and one or more auxiliary data units ((high-resolution images, image format) [Chen: col. 37, line 35]; (the original size of the image may be stored as metadata (e.g., height, width, and number of channels), and when the image is subsequently read from storage, the metadata can be checked to determine the actual dimensions of the image to avoid reading the empty characters or padding) [Chen: col. 37, line 29-34; Figs. 20, 71] ; (For example, visual data is commonly stored directly as files or in various types of databases (e.g., key-value, relational, and/or graph databases). Visual metadata is typically stored in databases, for example, while images and videos are typically stored as files. Moreover, different types of file systems and databases provide API functions in various programming and/or query languages in order to enable users to access and store data) [Chen: col. 32, line 67 – col. 33, line 7; Fig. 18]) as input to a neural network-based processor (FIGS. 24A-C illustrate an example embodiment of a multi-domain cascade convolutional neural network (CNN). FIGS. 25A-B, 26, 27, 28, 29, 30 and 31A-B illustrate the use of butterfly operations for a multi-domain convolutional neural network (CNN)) [Chen: col. 2, line 12-16; Figs. 24A-31C]; determining (determine) [Chen: col. 22, line 56] target features based at least on the one or more target data units (the converged node is capable of performing its described functionality while the data streams are in-flight versus post-facto. The converged node router performs an N-to-1 transformation, which may represent a range of processing capabilities, including but not limited to compression, encryption, transcoding, labeling, aggregation/grouping some flows into larger flows based on contextual commonality, sub-sampling, combination (e.g., stitching), and analytics ( e.g., which broadly refers to any type of analysis, whether it is statistical analysis, machine learning (ML), deep learning (DL) or some other form of artificial intelligence or machine learning)) [Chen: col. 90, line 44-56; Figs 58-64], and determining auxiliary features based at least on the one or more auxiliary data units ((the converged node is capable of performing its described functionality while the data streams are in-flight versus post-facto. The converged node router performs an N-to-1 transformation, which may represent a range of processing capabilities, including but not limited to compression, encryption, transcoding, labeling, aggregation/grouping some flows into larger flows based on contextual commonality, sub-sampling, combination (e.g., stitching), and analytics (e.g., which broadly refers to any type of analysis, whether it is statistical analysis, machine learning (ML), deep learning (DL) or some other form of artificial intelligence or machine learning)) [Chen: col. 90, line 44-56; Figs 58-64]; (the original size of the image may be stored as metadata (e.g., height, width, and number of channels), and when the image is subsequently read from storage, the metadata can be checked to determine the actual dimensions of the image to avoid reading the empty characters or padding) [Chen: col. 37, line 29-34; Figs. 20, 71]), by one or more first portions of the neural network-based processor (FIG. 8 illustrates an example embodiment of an architecture 800 for visual fog nodes. In some embodiments, for example, fog node architecture 800 may be used to implement the functionality of fog nodes 810 in a visual fog network or system (e.g., visual fog system 100 of FIG. 1). A fog node 810, for example, can include any node or component that ranges from the edge of a network to the cloud, inclusively. In the illustrated embodiment, fog node 810 includes various application programming interfaces (APis) that provide fundamental capabilities for fog node 810, such as auxiliary API 820, primitive visionAPI 830, and storageAPI 840. In some embodiments, for example, these APis may be used or implemented by lower-level algorithm developers. Auxiliary API 820 provides various fundamental functionality for fog node 810, such as security 822a, communication 822b, compression 822c ( e.g., codecs), and so forth. Primitive vision API 830 provides fundamental vision processing capabilities for fog node 810. For example, primitive vision API 830 provides access to a plurality of vision kernels 832 that can be used to perform primitive vision operations (e.g., person or object detection, facial recognition). Primitive vision API 830 may also provide access to various machine learning and/or neural network frameworks (e.g., Caffe, OpenCV, TensorFlow) ) [Chen: col. 21, line 12-36; Figs. 8-11]; combining the determined features or data derived from the determined features by a combination operation (In some embodiments, for example, the unified API could be used to retrieve and/or combine visual metadata and the original visual data from different storage locations. The unified API may also allow certain types of processing to be performed on visual data before it is returned to the requesting user. Further, the unified API may allow users to explicitly recognize visual entities such as images, feature vectors, and videos, and may simplify access to those visual entities based on their relationship with each other and with other entities associated with a particular vision application.) [Chen: col. 33, line 11-21; Fig. 18, 62]; and providing an output of the combination operation to a second portion of the neural network-based processor ((The flowchart then proceeds to block 8912 to determine if the processing associated with the visual data is complete (e.g., based on the output from the CNN(s) used in the current and/or preceding processing stages). For example, if the CNN in the current processing stage was unable to sufficiently interpret the visual data for purposes of deriving requisite information and/or reaching certain processing decision(s), the processing associated with the visual data may be incomplete. Accordingly, the flowchart proceeds back to block 8906, where the compressed data is transmitted to other processing device(s) in the network to perform additional stages of processing using different CNNs. The flowchart repeats in this manner as the compressed visual data is transmitted across the respective processing devices from the edge to the cloud, until it is eventually determined at block 8912 that the processing is complete) [Chen: col. 45, line 48-64; Fig. 89]; (routing the new data stream to its next intended upstream destination, which may be the northern direction in which the data was originally flowing, but may also include a broader dissemination, such as in the East-West direction to peer clouds or in the southern direction to interested parties) [Chen: col. 90, line 22-27; Figs. 62, 89]; (FIGS. 24A-C and FIG. 89 illustrate example embodiments of a multi-domain cascade convolutional neural network (CNN). In distributed visual analytics systems, for example, image and video is often compressed before transmission (e.g., from the pixel domain to a compressed domain), and subsequently decompressed after transmission (e.g., back to the pixel domain) before any processing can be performed, such as deep learning using neural networks. As an example, image and video captured by edge devices may be compressed and transmitted to the cloud, and then decompressed by the cloud before any further processing begins) [Chen: col. 40, line 63 – col. 41, line 8; Figs. 24A-C, 62, 89]). Regarding claim 17 and 25, Chen meets the claim limitations as set forth in claims 16 and 24. Chen further meets the claim limitations as follow. wherein the neural network-based processor (neural network processors (e.g., Movidius )) [col. 92, line 22-23; Figs. 66-68] is one of the following: neural network based in-loop filter; neural network-based post-processing filter (After the workloads finish executing, the distributed UVF execution framework 1712 generates an output 1713 resulting from execution of the UVFs 1709. For example, the output 1713 may include, or may be derived from, a filtered stream of visual data and/or metadata 1707 generated by execution of the UVFs 1709) [Chen: col. 31, line 61-66]; end-to-end learned neural network-based decoder (Code 604, which may be one or more instructions to be executed by processor 600, may be stored in memory 602, or may be stored in software, hardware, firmware, or any suitable combination thereof, or in any other internal or external component, device, element, or object where appropriate and based on particular needs. In one example, processor 600 can follow a program sequence of instructions indicated by code 604. Each instruction enters a front-end logic 606 and is processed by one or more decoders 608. The decoder may generate, as its output, a micro operation such as a fixed width micro operation in a predefined format, or may generate other instructions, microinstructions, or control signals that reflect the original code instruction. Front-end logic 606 may also include register renaming logic and scheduling logic, which generally allocate resources and queue the operation corresponding to the instruction for execution.) [Chen: col. 92, line 22-23; Figs. 66-68]; (The data aggregators 326 may collect data from any number of the sensors 328, and perform the back-end processing function for the analysis) [Chen: col. 12, line 21-24; Fig. 6]; or one of the neural networks in an end-to-end learned neural network-based decoder. Regarding claim 18 and 26, Chen meets the claim limitations as set forth in claims 16 and 24. Chen further meets the claim limitations as follow. motion-compensating the one or more data units (Moreover, CNN 2400 is designed to process compressed visual data 2402 as input (e.g., video sequence data compressed with a motion-compensated predictive coding scheme such as H.264)) [Chen: col. 41, line 64-67], or features extracted from the one or more data units (Deep learning neural networks, such as CNN s, are frequently used for image processing, including object/edge detection, segmentation, and classification, among other examples) [Chen: col. 35, line 29-32]; and using the motion-compensated data units as the one or more auxiliary data units (In some embodiments, for example, compressed visual data 2402 provided as input to CNN 2400 may first be partially decoded to separate and extract different syntax elements (e.g., motion vectors, macroblock (MB) coding modes, quantized prediction residuals), thus producing a subset of partial compression data 2404) [Chen: col. 42, line 1-6] or using the motion-compensated features as the auxiliary features (In some embodiments, for example, compressed visual data 2402 provided as input to CNN 2400 may first be partially decoded to separate and extract different syntax elements (e.g., motion vectors, macroblock (MB) coding modes, quantized prediction residuals), thus producing a subset of partial compression data 2404) [Chen: col. 42, line 1-6]. Regarding claim 19 and 27, Chen meets the claim limitations as set forth in claims 16 and 24. Chen further meets the claim limitations as follow. wherein the one or more target data units (FIG. 17 illustrates an example embodiment of a runtime processing pipeline 1700 for a visual fog architecture. In the illustrated embodiment, for example, a raw stream of visual data 1701 (e.g., video or images) captured by cameras or visual sensors in a visual fog architecture is provided as input to a stream ingress framework 1702) [Chen: col. 31, line 33-38; Fig. 17]) are one or more of the following: an image (images) [Chen: col. 7, line 54; Figs. 20, 71], a decompressed image (decompressed into the original image) [Chen: col. 38, line 16-17; Fig. 20], a video frame (video) [Chen: col. 7, line 54; Figs. 20, 71], or a decompressed video frame (decompressed into the original image) [Chen: col. 38, line 16-17; Fig. 20]; and wherein the one or more auxiliary data units (the original size of the image may be stored as metadata (e.g., height, width, and number of channels), and when the image is subsequently read from storage, the metadata can be checked to determine the actual dimensions of the image to avoid reading the empty characters or padding) [Chen: col. 37, line 29-34; Figs. 20, 71] are one of the following: other images (high-resolution images) [Chen: col. 37, line 35], other decompressed images (decompressed into the original image) [Chen: col. 38, line 16-17; Fig. 20], other video frames (video frames that contain similar content) [Chen: col. 43, line 46-47; Fig. 20], or other decompressed video frames (decompressed into the original image) [Chen: col. 38, line 16-17; Fig. 20]. Regarding claim 20 and 28, Chen meets the claim limitations as set forth in claims 16 and 24. Chen further meets the claim limitations as follow. wherein the apparatus (an apparatus) [Chen: Abstract; col. 112, line 25] is further caused to perform (Processor 600 can execute any type of instructions associated with algorithms, processes, or operations detailed herein. Generally, processor 600 can transform an element or an article (e.g., data) from one state or thing to another state or thing) [Chen: col. 18, line 41-45]: selecting the one or more auxiliary data units or auxiliary features based at least on one or more selection criteria, wherein the selection criteria are based at least on one of the following: a quality of the auxiliary data units; a resolution of the auxiliary data units (the original size of the image may be stored as metadata (e.g., height, width, and number of channels), and when the image is subsequently read from storage, the metadata can be checked to determine the actual dimensions of the image to avoid reading the empty characters or padding) [Chen: col. 37, line 29-34; Figs. 20, 71]; location of the auxiliary data units on a certain temporal layer; role of the auxiliary data units as a reference data; content of the auxiliary data units; or a distance of the one or more auxiliary data units compared to the one or more target data units. Regarding claim 21 and 29, Chen meets the claim limitations as set forth in claims 20 and 28. Chen further meets the claim limitations as follow. wherein to select one or more auxiliary data units (it requires one of the devices to select) [Chen: col. 68, line 34] based at least on one or more selection criteria comprises (selected from the database using feature extraction methodologies) [Chen: col. 59, line 17-18; col. 59, line 43-46]), and wherein the apparatus is further caused to select the one or more auxiliary data units that have a higher quality than the one or more target data units; or have a higher resolution than the one or more target data units ((high-resolution images, image format) [Chen: col. 37, line 35]; (the original size of the image may be stored as metadata (e.g., height, width, and number of channels), and when the image is subsequently read from storage, the metadata can be checked to determine the actual dimensions of the image to avoid reading the empty characters or padding) [Chen: col. 37, line 29-34; Figs. 20, 71]; (The dimensions and domain of the array depend on the resolution of the original image and therefore are calculated dynamically on a per-image basis. Since images often do not have an evenly divisible number of pixels in one or both dimensions, this occasionally results in the dimensions of an array not matching the original resolution of the image) [Chen: col. 38, line 58-63]; (For high-resolution images, image format 2200 improves the speed of common operations such as reading and writing, as well as the speed of operations used in image analytics, such as cropping and resizing.) [Chen: col. 37, line 35-38]); or are part of a lower temporal layer with respect to the one or more target data units ((To illustrate, existing deep learning CNNs (e.g., inception or ResNet CNN models) typically repeat an inner module multiple times, and the inner module aggregates the results from multiple convolution layers and/or the original input at the end (analogous to a bottleneck). For example, FIGS. 25A-B illustrate a traditional 27-layer inception model CNN 2500, and FIGS. 26 and 27 illustrate example inner modules 2600 and 2700 for an inception model CNN. In particular, FIG. 26 illustrates an inner module 2600 implemented without dimension reduction, while FIG. 27 illustrates an inner module 2700 implemented with dimension reduction) [Chen: col. 46, line 28-38; Figs. 2531]; (When a user creates an analytic image using VCL 5100, the analytic image schema is automatically set using the parameters described above in TABLE 1. VCL 5100 then creates a layer of abstraction with function calls of TileDB 5102 ( e.g., the array-database manager used in the illustrated embodiment) combined with specialized transformation operations to provide an interface to the analytic image) [Chen: col. 39, line 33-39]); or have similar content as the one or more target data units. Regarding claim 22 and 30, Chen meets the claim limitations as set forth in claims 16 and 24. Chen further meets the claim limitations as follow. wherein to select one or more auxiliary data units (it requires one of the devices to select) [Chen: col. 68, line 34] based at least on one or more selection criteria comprises (selected from the database using feature extraction methodologies) [Chen: col. 59, line 17-18; col. 59, line 43-46], and wherein the apparatus is further caused to select the one or more auxiliary data units (it requires one of the devices to select) [Chen: col. 68, line 34] by performing a hard or a soft attention operation (Existing visual processing systems, however, typically derive analytics and insights from raw images as input (e.g., by generating attention maps), which can compromise the privacy of people captured in the images, as it may reveal their identity and/or other personal information) [Chen: col. 99, line 10-14] based at least on one or more attention neural networks (For example, primitive vision API 830 provides access to a plurality of vision kernels 832 that can be used to perform primitive vision operations (e.g., person or object detection, facial recognition). Primitive vision API 830 may also provide access to various machine learning and/or neural network frameworks (e.g., Caffe, OpenCV, TensorFlow)) [Chen: col. 21, line 30-36], wherein the result of the attention operations functions as a selection criterion (There is an implicit mapping that can be taken into account by an analytics component. For instance, pages or sections visited online can be mapped to areas visited in the physical store. The resulting integrated customer model produces better results from analytics (FIG. 82, 8221) that can be used to improve the interactions between the business and the customer (e.g., by providing better product recommendations (FIG. 82, 8206, 8219, 8220, 8224)). The described solution pays particular attention to the interaction that the customer has with products while visiting the store, particularly for products that the customer does not end up buying (FIG. 82, 8215-8217, 8219)) [Chen: col. 106, line 16-30; Fig. 84]. Regarding claim 23, Chen meets the claim limitations as follows: An apparatus (an apparatus) [Chen: Abstract; col. 112, line 25] comprising: at least one processor (a processor 552, which may be a microprocessor, a multi-core processor, a multithreaded processor, an ultra-low voltage processor, an embedded processor, or other known processing element.) [Chen: col. 14, line 19-22; Figs. 5-7, 64]; and at least one memory (a system memory) [Chen: col. 14, line 41-42; Figs. 5-7, 64] storing instructions ((computer instructions) [Chen: col. 112, line 47]; (Program code, such as code 730 illustrated in FIG. 7, may 30 be applied to input instructions to perform the functions described herein and generate output information.) [Chen: col. 20, line 29-31; Fig. 7]) that, when executed by the at least one processor, cause the apparatus at least to perform ((Processor 600 can execute any type of instructions associated with algorithms, processes, or operations detailed herein. Generally, processor 600 can transform an element or an article (e.g., data) from one state or thing to another state or thing) [Chen: col. 18, line 41-45]; (In another example, some activities outlined herein may be implemented with fixed logic or programmable logic (for example, software and/or computer instructions executed by a processor) and the elements identified herein could be some type of a programmable processor, programmable digital logic (for example, a field programmable gate array (FPGA), an erasable programmable read only memory (EPROM), an electrically erasable programmable read only memory (EEPROM)), an ASIC that includes digital logic, software, code, electronic instructions, flash memory, optical disks, CD-ROMs, DVD ROMs, magnetic or optical cards, other types of machine-readable mediums suitable for storing electronic instructions, or any suitable combination thereof) [Chen: col. 112, line 45-58]): determining (determine) [Chen: col. 22, line 56] target features based at least on the one or more target data units (the converged node is capable of performing its described functionality while the data streams are in-flight versus post-facto. The converged node router performs an N-to-1 transformation, which may represent a range of processing capabilities, including but not limited to compression, encryption, transcoding, labeling, aggregation/grouping some flows into larger flows based on contextual commonality, sub-sampling, combination (e.g., stitching), and analytics ( e.g., which broadly refers to any type of analysis, whether it is statistical analysis, machine learning (ML), deep learning (DL) or some other form of artificial intelligence or machine learning)) [Chen: col. 90, line 44-56; Figs 58-64]; determining (determine) [Chen: col. 22, line 56] auxiliary features based at least on the one or more auxiliary data units (the converged node is capable of performing its described functionality while the data streams are in-flight versus post-facto. The converged node router performs an N-to-1 transformation, which may represent a range of processing capabilities, including but not limited to compression, encryption, transcoding, labeling, aggregation/grouping some flows into larger flows based on contextual commonality, sub-sampling, combination (e.g., stitching), and analytics ( e.g., which broadly refers to any type of analysis, whether it is statistical analysis, machine learning (ML), deep learning (DL) or some other form of artificial intelligence or machine learning)) [Chen: col. 90, line 44-56; Figs 58-64]; (the original size of the image may be stored as metadata (e.g., height, width, and number of channels), and when the image is subsequently read from storage, the metadata can be checked to determine the actual dimensions of the image to avoid reading the empty characters or padding) [Chen: col. 37, line 29-34; Figs. 20, 71]); determining (determine) [Chen: col. 22, line 56] a loss term (prediction error 6532) [Chen: col. 56, line 37] based at least on the target features and the auxiliary features (Further, while the slave macroblock 6512 and the matching master macroblock 6522 may be similar, it is unlikely that they are an exact match. Accordingly, the differences between the slave macro block 6512 and the matching master macroblock 6522 are computed in order to produce a prediction error 6532) [Chen: col. 56, line 32-37], or data derived from the target features and the auxiliary features (the converged node is capable of performing its described functionality while the data streams are in-flight versus post-facto. The converged node router performs an N-to-1 transformation, which may represent a range of processing capabilities, including but not limited to compression, encryption, transcoding, labeling, aggregation/grouping some flows into larger flows based on contextual commonality, sub-sampling, combination (e.g., stitching), and analytics ( e.g., which broadly refers to any type of analysis, whether it is statistical analysis, machine learning (ML), deep learning (DL) or some other form of artificial intelligence or machine learning)) [Chen: col. 90, line 44-56; Figs 58-64]; and training a neural network-based processor (training classification model) [Chen: col. 59, line 5; Fig. 40C]; (neural network processors (e.g., Movidius )) [col. 92, line 22-23; Figs. 66-68]) to minimize a loss function comprising the determined loss term (minimize: cTx (objective term) subject to: Ax < b (inequality constraint)) [Chen: col. 64, line 42-43]. Regarding claim 24, Chen meets the claim limitations, as follows: A method (methodologies) [Chen: Abstract; col. 9, line 10] comprising: receiving one or more target data units ((In the illustrated example, the original image is first received 2202) [Chen: col. 37, line 11-12; Figs. 13, 19, 22]; (FIG. 17 illustrates an example embodiment of a runtime processing pipeline 1700 for a visual fog architecture. In the illustrated embodiment, for example, a raw stream of visual data 1701 (e.g., video or images) captured by cameras or visual sensors in a visual fog architecture is provided as input to a stream ingress framework 1702) [Chen: col. 31, line 33-38; Fig. 17]) and one or more auxiliary data units ((high-resolution images, image format) [Chen: col. 37, line 35]; (the original size of the image may be stored as metadata (e.g., height, width, and number of channels), and when the image is subsequently read from storage, the metadata can be checked to determine the actual dimensions of the image to avoid reading the empty characters or padding) [Chen: col. 37, line 29-34; Figs. 20, 71]; (For example, visual data is commonly stored directly as files or in various types of databases (e.g., key-value, relational, and/or graph databases). Visual metadata is typically stored in databases, for example, while images and videos are typically stored as files. Moreover, different types of file systems and databases provide API functions in various programming and/or query languages in order to enable users to access and store data) [Chen: col. 32, line 67 – col. 33, line 7; Fig. 18]) as input to a neural network-based processor (FIGS. 24A-C illustrate an example embodiment of a multi-domain cascade convolutional neural network (CNN). FIGS. 25A-B, 26, 27, 28, 29, 30 and 31A-B illustrate the use of butterfly operations for a multi-domain convolutional neural network (CNN)) [Chen: col. 2, line 12-16; Figs. 24A-31C]; determining (determine) [Chen: col. 22, line 56] target features based at least on the one or more target data units (the converged node is capable of performing its described functionality while the data streams are in-flight versus post-facto. The converged node router performs an N-to-1 transformation, which may represent a range of processing capabilities, including but not limited to compression, encryption, transcoding, labeling, aggregation/grouping some flows into larger flows based on contextual commonality, sub-sampling, combination (e.g., stitching), and analytics ( e.g., which broadly refers to any type of analysis, whether it is statistical analysis, machine learning (ML), deep learning (DL) or some other form of artificial intelligence or machine learning)) [Chen: col. 90, line 44-56; Figs 58-64], and determining auxiliary features based at least on the one or more auxiliary data units (the converged node is capable of performing its described functionality while the data streams are in-flight versus post-facto. The converged node router performs an N-to-1 transformation, which may represent a range of processing capabilities, including but not limited to compression, encryption, transcoding, labeling, aggregation/grouping some flows into larger flows based on contextual commonality, sub-sampling, combination (e.g., stitching), and analytics ( e.g., which broadly refers to any type of analysis, whether it is statistical analysis, machine learning (ML), deep learning (DL) or some other form of artificial intelligence or machine learning)) [Chen: col. 90, line 44-56; Figs 58-64]; (the original size of the image may be stored as metadata (e.g., height, width, and number of channels), and when the image is subsequently read from storage, the metadata can be checked to determine the actual dimensions of the image to avoid reading the empty characters or padding) [Chen: col. 37, line 29-34; Figs. 20, 71], by one or more first portions of the neural network-based processor (FIG. 8 illustrates an example embodiment of an architecture 800 for visual fog nodes. In some embodiments, for example, fog node architecture 800 may be used to implement the functionality of fog nodes 810 in a visual fog network or system (e.g., visual fog system 100 of FIG. 1). A fog node 810, for example, can include any node or component that ranges from the edge of a network to the cloud, inclusively. In the illustrated embodiment, fog node 810 includes various application programming interfaces (APis) that provide fundamental capabilities for fog node 810, such as auxiliary API 820, primitive visionAPI 830, and storageAPI 840. In some embodiments, for example, these APis may be used or implemented by lower-level algorithm developers. Auxiliary API 820 provides various fundamental functionality for fog node 810, such as security 822a, communication 822b, compression 822c ( e.g., codecs), and so forth. Primitive vision API 830 provides fundamental vision processing capabilities for fog node 810. For example, primitive vision API 830 provides access to a plurality of vision kernels 832 that can be used to perform primitive vision operations (e.g., person or object detection, facial recognition). Primitive vision API 830 may also provide access to various machine learning and/or neural network frameworks (e.g., Caffe, OpenCV, TensorFlow) ) [Chen: col. 21, line 12-36; Figs. 8-11]; combining the determined features or data derived from the determined features by a combination operation (In some embodiments, for example, the unified API could be used to retrieve and/or combine visual metadata and the original visual data from different storage locations. The unified API may also allow certain types of processing to be performed on visual data before it is returned to the requesting user. Further, the unified API may allow users to explicitly recognize visual entities such as images, feature vectors, and videos, and may simplify access to those visual entities based on their relationship with each other and with other entities associated with a particular vision application.) [Chen: col. 33, line 11-21; Fig. 18, 62]; and providing an output of the combination operation to a second portion of the neural network-based processor ((The flowchart then proceeds to block 8912 to determine if the processing associated with the visual data is complete (e.g., based on the output from the CNN(s) used in the current and/or preceding processing stages). For example, if the CNN in the current processing stage was unable to sufficiently interpret the visual data for purposes of deriving requisite information and/or reaching certain processing decision(s), the processing associated with the visual data may be incomplete. Accordingly, the flowchart proceeds back to block 8906, where the compressed data is transmitted to other processing device(s) in the network to perform additional stages of processing using different CNNs. The flowchart repeats in this manner as the compressed visual data is transmitted across the respective processing devices from the edge to the cloud, until it is eventually determined at block 8912 that the processing is complete) [Chen: col. 45, line 48-64; Fig. 89]; (routing the new data stream to its next intended upstream destination, which may be the northern direction in which the data was originally flowing, but may also include a broader dissemination, such as in the East-West direction to peer clouds or in the southern direction to interested parties) [Chen: col. 90, line 22-27; Figs. 62, 89]; (FIGS. 24A-C and FIG. 89 illustrate example embodiments of a multi-domain cascade convolutional neural network (CNN). In distributed visual analytics systems, for example, image and video is often compressed before transmission (e.g., from the pixel domain to a compressed domain), and subsequently decompressed after transmission (e.g., back to the pixel domain) before any processing can be performed, such as deep learning using neural networks. As an example, image and video captured by edge devices may be compressed and transmitted to the cloud, and then decompressed by the cloud before any further processing begins) [Chen: col. 40, line 63 – col. 41, line 8; Figs. 24A-C, 62, 89]). Reference Notice Additional prior arts, included in the Notice of Reference Cited, made of record and not relied upon is considered pertinent to applicant's disclosure. Contact Information Any inquiry concerning this communication or earlier communications from the examiner should be directed to Philip Dang whose telephone number is (408) 918-7529. The examiner can normally be reached on Monday-Thursday between 8:30 am - 5:00 pm (PST). Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Sath Perungavoor can be reached on 571-272-7455. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of an application may be obtained from the Patent Application Information Retrieval (PAIR) system. Status information for published applications may be obtained from either Private PAIR or Public PAIR. Status information for unpublished applications is available through Private PAIR only. For more information about the PAIR system, see http://pair-direct.uspto.gov. Should you have questions on access to the Private PAIR system, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative or access to the automated information system, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000./Philip P. Dang/Primary Examiner, Art Unit 2488
Read full office action

Prosecution Timeline

Dec 28, 2024
Application Filed
Jul 17, 2026
Non-Final Rejection mailed — §102, §112 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12707047
IMAGE ENCODING DEVICE, IMAGE ENCODING METHOD, AND IMAGE ENCODING PROGRAM, AND IMAGE DECODING DEVICE, IMAGE DECODING METHOD, AND IMAGE DECODING PROGRAM
1y 8m to grant Granted Aug 11, 2026
Patent 12701231
QUANTIZATION OF RESIDUALS IN VIDEO CODING
4y 7m to grant Granted Aug 04, 2026
Patent 12689733
VIDEO ENCODING METHOD AND DEVICE, AND VIDEO DECODING METHOD AND DEVICE
2y 0m to grant Granted Jul 21, 2026
Patent 12689770
SCALING LIST-BASED VIDEO OR IMAGE CODING
1y 6m to grant Granted Jul 21, 2026
Patent 12682475
LARGE DEPTH-OF-FIELD MICROSCOPIC STRUCTURED-LIGHT 3D IMAGING
2y 4m to grant Granted Jul 14, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
78%
Grant Probability
99%
With Interview (+30.4%)
2y 7m (~1y 0m remaining)
Median Time to Grant
Low
PTA Risk
Based on 496 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month