Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
DETAILED ACTION
The United States Patent & Trademark Office appreciates the application that is submitted by the inventor/assignee. The United States Patent & Trademark Office reviewed the following application and has made the following comments below.
Priority
This application claims benefit of foreign priority under 35 U.S.C. 119(a)-(d) of CN 202210753534.6, filed in China on 06/28/2022, and PCT/KR2023/008963, filed in Korea on 06/27/2023.
Information Disclosure Statement
The information disclosure statements (IDS) submitted on 12/30/2024, 09/29/2025, and 04/30/2026 are in compliance with the provisions of 37 CFR 1.97. Accordingly, the information disclosure statements are being considered by the examiner.
Specification
The disclosure is objected to because of the following informalities:
In paragraph [0004], on line 1, “a an image processing method” is suggested to read “an image processing method”.
In paragraph [0004], on line 2, “storage medium be capable” is suggested to read “storage medium capable”.
In paragraph [0061], on line 2, “will grow in a quadratic rate” is suggested to read “will grow at a quadratic rate”.
In paragraph [0062], on line 3, “Regarding to the model” is suggested to read “Regarding the model”.
In paragraph [0094], on line 9, “that is more deserving” is suggested to read “that are more deserving”.
In paragraph [0097], on line 1, “(408)by” is suggested to read “(408) by”. It appears a space is missing between “(408)” and “by”.
In paragraph [0098], on line 2, “kernel that enable extraction” is suggested to read “kernel that enables extraction”.
In paragraph [0106], on line 6, “(in the first neural network In the first neural network” is suggested to read “(in the first neural network”.
In paragraph [0110], on line 7, “as an example:” is suggested to read “as an example.” It appears that the sentence ending in “example” has a colon at the end of the sentence instead of a period.
In paragraph [0111], on line 1, “cross-grain attention module” is suggested to read “cross-granularity attention module”.
In paragraph [0122], on lines 3-4, “performing at least one down sampling for the first image patches” is suggested to read “performing at least one down sampling method (or operation) for the first image patches”.
In paragraph [0124], on line 1, “each down sampling may” is suggested to read “each down sampling method (or operation) may”.
In paragraph [0126], on line 1, “each down sampling may” is suggested to read “each down sampling method (or operation) may”.
In paragraph [0133], on lines 8-9, “by a predetermined by a predetermined multiple” is suggested to read “by a predetermined multiple”.
In paragraph [0166], on line 1, “embodimentmay” is suggested to read “embodiment may”. It appears there is a space missing between “embodiment” and “may”.
In paragraph [0214], on line 3, “one frame of a video frame” is suggested to read “one frame of a video”.
In paragraph [0329], on line 1, “the terms ‘first’, ‘second’, ‘third’, ‘fourth’, ‘third’ and ‘fourth’” is suggested to read “the terms ‘first’, ‘second’, ‘third’, ‘fourth’, ‘fifth’ and ‘sixth’”. It appears “third” and “fourth” were repeated twice.
Appropriate correction is required.
Drawings
The drawings are objected to under 37 CFR 1.83(a) because they fail to show “image patch A shown in the upper left corner of FIG. 6 as an example” and “taking image patches A-D shown in the upper right corner of Fig. 6 as an example” as described in the specification in Paragraphs [0109] and [0110]. Fig. 6 of applicant’s disclosure is shown below. It is unclear what is being referred to as “image patch A shown in the upper left corner of FIG. 6” and “image patches A-D shown in the upper right corner of Fig. 6” since the figure (Fig. 6) does not depict image patches labeled A-D.
PNG
media_image1.png
646
805
media_image1.png
Greyscale
Any structural detail that is essential for a proper understanding of the disclosed invention should be shown in the drawing. MPEP § 608.02(d). Corrected drawing sheets in compliance with 37 CFR 1.121(d) are required in reply to the Office action to avoid abandonment of the application. Any amended replacement drawing sheet should include all of the figures appearing on the immediate prior version of the sheet, even if only one figure is being amended. The figure or figure number of an amended drawing should not be labeled as “amended.” If a drawing figure is to be canceled, the appropriate figure must be removed from the replacement sheet, and where necessary, the remaining figures must be renumbered and appropriate changes made to the brief description of the several views of the drawings for consistency. Additional replacement sheets may be necessary to show the renumbering of the remaining figures. Each drawing sheet submitted after the filing date of an application must be labeled in the top margin as either “Replacement Sheet” or “New Sheet” pursuant to 37 CFR 1.121(d). If the changes are not accepted by the examiner, the applicant will be notified and informed of any required corrective action in the next Office action. The objection to the drawings will not be held in abeyance.
Claim Interpretation
The following is a quotation of 35 U.S.C. 112(f):
(f) Element in Claim for a Combination. – An element in a claim for a combination may be expressed as a means or step for performing a specified function without the recital of structure, material, or acts in support thereof, and such claim shall be construed to cover the corresponding structure, material, or acts described in the specification and equivalents thereof.
The following is a quotation of pre-AIA 35 U.S.C. 112, sixth paragraph:
An element in a claim for a combination may be expressed as a means or step for performing a specified function without the recital of structure, material, or acts in support thereof, and such claim shall be construed to cover the corresponding structure, material, or acts described in the specification and equivalents thereof.
The claims in this application are given their broadest reasonable interpretation using the plain meaning of the claim language in light of the specification as it would be understood by one of ordinary skill in the art. The broadest reasonable interpretation of a claim element (also commonly referred to as a claim limitation) is limited by the description in the specification when 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is invoked.
As explained in MPEP § 2181, subsection I, claim limitations that meet the following three-prong test will be interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph:
(A) the claim limitation uses the term “means” or “step” or a term used as a substitute for “means” that is a generic placeholder (also called a nonce term or a non-structural term having no specific structural meaning) for performing the claimed function;
(B) the term “means” or “step” or the generic placeholder is modified by functional language, typically, but not always linked by the transition word “for” (e.g., “means for”) or another linking word or phrase, such as “configured to” or “so that”; and
(C) the term “means” or “step” or the generic placeholder is not modified by sufficient structure, material, or acts for performing the claimed function.
Use of the word “means” (or “step”) in a claim with functional language creates a rebuttable presumption that the claim limitation is to be treated in accordance with 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph. The presumption that the claim limitation is interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is rebutted when the claim limitation recites sufficient structure, material, or acts to entirely perform the recited function.
Absence of the word “means” (or “step”) in a claim creates a rebuttable presumption that the claim limitation is not to be treated in accordance with 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph. The presumption that the claim limitation is not interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is rebutted when the claim limitation recites function without reciting sufficient structure, material or acts to entirely perform the recited function.
Claim limitations in this application that use the word “means” (or “step”) are being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, except as otherwise indicated in an Office action. Conversely, claim limitations in this application that do not use the word “means” (or “step”) are not being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, except as otherwise indicated in an Office action.
This application includes one or more claim limitations that do not use the word “means,” but are nonetheless being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, because the claim limitation(s) uses a generic placeholder that is coupled with functional language without reciting sufficient structure to perform the recited function and the generic placeholder is not preceded by a structural modifier.
Such claim limitation(s) is/are:
“a first obtaining module” in claim 15
“a first processing module” in claims 15, 16, 17, and 18
“a first recognition module” in claims 15, 16, 17, and 18
Because this/these claim limitation(s) is/are being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, it/they is/are being interpreted to cover the corresponding structure described in the specification as performing the claimed function, and equivalents thereof.
If applicant does not intend to have this/these limitation(s) interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, applicant may: (1) amend the claim limitation(s) to avoid it/them being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph (e.g., by reciting sufficient structure to perform the claimed function); or (2) present a sufficient showing that the claim limitation(s) recite(s) sufficient structure to perform the claimed function so as to avoid it/them being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph.
For the sake of further prosecution, the Examiner will treat “a first obtaining module”, “a first processing module” and “a first recognition module” all as hardware or software configured to perform their respective functions/operations.
Claim Objections
Claims 5, 10, and 19 are objected to because of the following informalities:
In Claim 5, on lines 14 and 15, “performing at least one down sampling” is suggested to read “performing at least one down sampling method” or “performing at least one down sampling operation”.
In Claim 5, on line 15, “at least one down sampling” is suggested to read “at least one down sampling method” or “at least one down sampling operation”.
In Claim 10, on line 7, “one or more highlight” is suggested to read “one or more highlights”.
In Claim 19, on lines 4-5, “wherein the at least one processor is further configured to: for determining the recognition result” is suggested to read “wherein the at least one processor is further configured to: determine the recognition result”.
In Claim 19, on line 10, “for determining the recognition result” is suggested to read “determine the recognition result”.
In Claim 19, on line 14, “for obtaining the fourth image patches” is suggested to read “obtain the fourth image patches”.
In Claim 19, on line 17, “perform at least one down sampling” is suggested to read “perform at least one down sampling method” or “perform at least one down sampling operation”.
In Claim 19, on line 18, “performing the at least one down sampling” is suggested to read “performing the at least one down sampling method” or “performing the at least one down operation”.
Appropriate correction is required.
Claim Rejections - 35 USC § 112
The following is a quotation of 35 U.S.C. 112(b):
(B) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention.
The following is a quotation of 35 U.S.C. 112 (pre-AIA ), second paragraph:
The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the applicant regards as his invention.
Claims 15-18 are rejected under 35 U.S.C. 112(b) or 35 U.S.C. 112 (pre-AIA ), second paragraph, as being indefinite for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor, or for pre-AIA the applicant regards as the invention. The Examiner strongly suggested that appropriate corrections be made to clarify the claim scope.
Claim limitations:
“a first obtaining module” (Claim 15)
“a first processing module” (Claims 15, 16, 17, and 18)
“a first recognition module” (Claims 15, 16, 17, and 18)
invoke 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph. However, the written
description fails to disclose the corresponding structure, material, or acts for performing the entire claimed function and to clearly link the structure, material, or acts to the function.
The “first obtaining module” is mentioned in multiple paragraphs, for example, paragraphs [0020], [0263], and [0264] of the specification, as well as in Figure 18 (reference character 1801) of the drawings. However, the specification only discloses the claimed function of the first obtaining module. For example, paragraph [0264] of Applicant’s specification states, “the first obtaining module 1801 may be configured to obtain first image patches corresponding to the image to be processed”. In addition, Fig. 18 is a schematic diagram and does not disclose the corresponding structure of the “first obtaining module”.
Next, the “first processing module” is mentioned in several paragraphs throughout the specification, specifically paragraphs [0020], [0022], [0024], [0263], [0265], [0267], [0269], [0271], [0273], and [0275-277], as well as in Figure 18 (reference character 1802) of the drawings. Paragraph [0234] of applicant’s specification states, “the first processing module 1802 may be configured to divide the first image patches into at least two groups via a window self-attention network; and to determine the attention information among first image patches in each group of first image patches, respectively for each group of the first image patches; and to obtain second image patches comprising local attention information”, which recites only the function of the first processing module.
Lastly, the “first recognition module” is mentioned in paragraphs [0020-24], [0263], [0266], [0268], [0270], [0272], [0274], [0280], [0283], and [0287] as well as Figure 18 (reference character 1803) of the drawings. However, the specification only discloses the claimed function of the “first recognition module”. For example, applicant’s specification states, “the first recognition module 1803 may be configured to determine the recognition result of the image to be processed, based on the second image patches.”
This disclosure is not sufficient because it fails to disclose the structure of the above elements. Therefore, the claims are indefinite and are rejected under 35 U.S.C. 112(b) or pre-AIA 35 U.S.C. 112, second paragraph.
Applicant may:
Amend the claim so that the claim limitation will no longer be interpreted as a limitation under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph;
Amend the written description of the specification such that it expressly recites what structure, material, or acts perform the entire claimed function, without introducing any new matter (35 U.S.C. 132(a)); or
Amend the written description of the specification such that it clearly links the structure, material, or acts disclosed therein to the function recited in the claim, without introducing any new matter (35 U.S.C. 132(a)).
If applicant is of the opinion that the written description of the specification already implicitly or inherently discloses the corresponding structure, material, or acts and clearly links them to the function so that one of ordinary skill in the art would recognize what structure, material, or acts perform the claimed function, applicant should clarify the record by either:
Amending the written description of the specification such that it expressly recites the corresponding structure, material, or acts for performing the claimed function and clearly links or associates the structure, material, or acts to the claimed function, without introducing any new matter (35 U.S.C. 132(a)); or
Stating on the record what the corresponding structure, material, or acts, which are implicitly or inherently set forth in the written description of the specification, perform the claimed function. For more information, see 37 CFR 1.75(d) and MPEP §§ 608.01(o) and 2181.
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103(a) which forms the basis for all obviousness rejections set forth in this Office action:
(a) A patent may not be obtained though the invention is not identically disclosed or described as set forth in section 102 of this title, if the differences between the subject matter sought to be patented and the prior art are such that the subject matter as a whole would have been obvious at the time the invention was made to a person having ordinary skill in the art to which said subject matter pertains. Patentability shall not be negatived by the manner in which the invention was made.
The factual inquiries set forth in Graham v. John Deere Co., 383 U.S. 1, 148 USPQ 459 (1966), that are applied for establishing a background for determining obviousness under 35 U.S.C. 103(a) are summarized as follows:
1. Determining the scope and contents of the prior art.
2. Ascertaining the differences between the prior art and the claims at issue.
3. Resolving the level of ordinary skill in the pertinent art.
4. Considering objective evidence present in the application indicating obviousness or nonobviousness.
Claims 1-3, 10-11, 13-17, and 20 are rejected under 35 U.S.C. 103(a) as being unpatentable over Vaswani et al. (U.S. Patent Pub. No. 20210390410 A1, hereafter referred to as Vaswani) in view of Hatamizadeh et al. (U.S. Patent Pub No. 20230394781 A1, hereafter referred to as Hatamizadeh).
Regarding Claim 1, Vaswani teaches the following:
An image processing method (Abstract, Vaswani teaches a method for processing images using a computer vision neural network that has one or more local self-attention layers.)
comprising:
obtaining first image patches (Paragraphs [0033-34], Fig. 2, Vaswani teaches the layer input (210) is a [4, 4, c] “image”, where [4, 4] represents the height and width dimensions and c represents the channel dimension. While Fig. 2 shows the layer input (210) being a single “image,” i.e., a single tensor generated from a single input image, in some implementations, each layer input includes a “batch” dimension, where the neural network processes a batch of multiple input images in parallel and the layer input (210) includes a respective index along the batch dimension for each input image in the batch.) corresponding to an image to be processed (Paragraph [0034], Vaswani teaches the layer input (210) being a single “image” i.e., a single tensor generated from a single input image.);
dividing the first image patches into at least two groups via a window self-attention network (Paragraphs [0036-38], Fig. 2, Vaswani teaches the local self-attention layer can group the elements into different query blocks in the (height, width) domains. That is, for every element corresponding to a particular height index and a particular width index, the local self-attention layer can assign the element to a particular query block. The blocking performed by the layer divides the block input (210) into (H/b x W/b) non-overlapping (b, b, c) blocks, where b is a block size value for a layer. In Fig. 2, b is equal to 2 and the system has divided the input block (210) into a “blocked” image that includes four 2 element by 2 element query blocks (220).);
PNG
media_image2.png
223
314
media_image2.png
Greyscale
obtaining second image patches comprising local attention information (Paragraph [0043], Fig. 2, Vaswami teaches the local self-attention layer can generate a block attention output (250) for the given query block (220) that includes a respective attention output for each element in the query block. Because the attention is “local” within the query block, the layer can generate the block attention outputs (250) for all of the query blocks (220) in parallel.);
PNG
media_image3.png
376
901
media_image3.png
Greyscale
and determining a recognition result of the image to be processed (Paragraphs [0017-20], Vaswani teaches processing an input that includes an image to generate a corresponding output, e.g., a classification output, a regression output, or a combination thereof, for the computer vision task. As an example, the neural network can be configured to process an image to generate a classification output that includes a respective score corresponding to each of multiple categories. The categories may be classes of objects (e.g., dog, cat, person, and the like), and the image may belong to the category if it depicts an object included in the object class according to the category. Other examples include generating a pixel-level classification output and a regression output (such as coordinates of a bounding box enclosing the object).) based on the second image patches (Paragraphs [0069], [0017], [0022], [0025], Vaswani teaches a block attention output (250 in Fig. 2), comprising determining a respective query for each element in the query block (220), determining a respective key for each element of the corresponding context block (240), determining a value for each element of the corresponding context block, and using the determined query, keys, and values to generate a respective attention output for each element of the query block. The feature representation of the input image can be one or more tensor numeric values that represent learned properties of the input image such as a single feature map. The output neural network maps the feature representation to an appropriate output for the computer vision task. The output for the computer vision task is one or more of a predicted classification, semantic segmentation, or an object detection for the input image.).
Vaswani does not explicitly disclose the following:
determining global attention information among the first image patches in the at least two groups of the first image patches, respectively.
Hatamizadeh is in the same field of art of extracting local features from local windows/patches of an image. Further, Hatamizadeh teaches determining global attention information among the first image patches (Paragraphs [0059], [0070], [0033], Fig. 5B, Hatamizadeh teaches global attention is computed jointly with local attention. A global self-attention module that accesses, per local window of the plurality of local windows within the input image, global features extracted from an entirety of an input image, or from at least a portion of the input image outside the local window. The global features may be of any defined category (e.g., textures, shape descriptors, etc.).) in the at least two groups (Paragraph [0029], Hatamizadeh teaches the input image is apportioned into a plurality of local windows. The Examiner interprets a “local window” to be a group of image patches since the local window includes a plurality of image patches.) of the first image patches, respectively (Paragraph [0029], Hatamizadeh teaches each of the local windows includes a plurality of image patches, which may be blocks or other image portions each composed of one or more pixels or other image elements.).
PNG
media_image4.png
298
699
media_image4.png
Greyscale
Therefore, it would have been obvious to one having ordinary skill in the art before the effective filing date of the claimed invention to modify the invention of Vaswani by computing global attention information/extracting global features for each window within the input image that is taught by Hatamizadeh, to make the invention that computes both local and global attention information for the image patches of the image; thus, one of ordinary skilled in the art would be motivated to combine the references since there is a need for vision transformers to be able to capture long-range spatial dependencies in a less computationally expensive manner (Hatamizadeh, Paragraph [0005]). In addition, by computing both local and global self-attention during image processing, both short-range and long-range spatial dependencies may be respectively modeled by the vision transformer, which improves the quality of the feature representations obtained by the vision transformer.
Thus, the claimed subject matter would have been obvious to a person having ordinary skill in the art before the effective filing date of the claimed invention.
In regards to Claim 2, Vaswani in view of Hatamizadeh discloses wherein the determining the recognition result of the image to be processed based on the second image patches comprises:
determining, by a global token generator (Paragraphs [0016], [0034], [0052], Fig. 6, Hatamizadeh teaches a global token generator.),
PNG
media_image5.png
346
701
media_image5.png
Greyscale
at least one global token corresponding to at least part of the image to be processed (Paragraphs [0034-35], [0061-63], Fig. 8, Hatamizadeh teaches global features may be extracted from the entirety of the input image by a global token generator of the vision transformer. The global features may be used as a global query token.),
PNG
media_image6.png
429
596
media_image6.png
Greyscale
based on the first image patches (Paragraph [0008], Hatamizadeh teaches a feature map is generated for an image. The feature map is processed to generate global query tokens that spatially correspond with local tokens of each window of a plurality of local windows within an image.); and
determining the recognition result of the image to be processed (Paragraph [0024], Hatamizadeh teaches the downstream task can then process the given input (input embeddings) to provide for example, image classification, object detection, instance segmentation, semantic segmentation. The Examiner interprets performing object detection on an image determines a “recognition result” i.e., recognizes an object in the image.), based on the at least one global token and the second image patches (Paragraphs [0024], [0037], Hatamizadeh teaches the vision transformer that is configured to process images, using both local and global self-attention, to derive information from those images. The information derived by the vision transformer may be feature representations for an input image. The derived information (such as long-range (global) and short-range (local) dependencies) may then be provided as input embeddings, to a computer vision related downstream task.).
In regards to Claim 3, Vaswani in view of Hatamizadeh discloses the image processing method of claim 2, wherein the global token generator (Paragraphs [0061-64], Fig. 6, Hatamizadeh teaches a global token generator.) comprises a kernel generator (Paragraphs [0050-51], Vaswani teaches applying a 3D convolution that has a kernel made up of ones and zeros. The Examiner interprets a “kernel generator” to be a convolutional layer in light of Applicant’s specification, which states, “the kernel generator may employ, but is not limited to, convolutional layers, fully connecting layers, etc. It should be understood that, for some cases, it may be equivalent to use convolution layer or a fully connecting layer (Paragraph [0091]).”), and
wherein the determining, by the global token generator (Paragraphs [0061-64], Hatamizadeh teaches a global token generator.), the at least one global token (Paragraph [0063], Hatamizadeh teaches generating global query tokens.) corresponding to the at least part of the image to be processed based on the first image patches (Paragraph [0063], Hatamizadeh teaches the global query tokens encompass information across the entire input feature map for an input image.), comprises:
generating at least one kernel for the image to be processed (Paragraphs [0050-51], Vaswani teaches performing convolution using a kernel; that includes a ‘1’ at each location corresponding to an element in the given context block.), by the kernel generator (Paragraphs [0050-51], Vaswani teaches applying a 3D convolution that has a kernel made up of ones and zeros. The Examiner interprets a “kernel generator” to be a convolutional layer in light of Applicant’s specification, which states, “the kernel generator may employ, but is not limited to, convolutional layers, fully connecting layers, etc. It should be understood that, for some cases, it may be equivalent to use convolution layer or a fully connecting layer (Paragraph [0091]).”); and
determining the at least one global token (Paragraphs [0061- 63], Fig. 6, Hatamizadeh teaches generating global query tokens.) respectively corresponding to the at least one kernel (Paragraphs [0063], [0074] Hatamizadeh teaches global query tokens for interaction with local key and value features per local window when computing global self-attention.), based on the at least one kernel (Paragraphs [0063], [0074], Hatamizadeh teaches global query tokens that spatially correspond with local tokens of each local window of a plurality of local windows within the image, such that the local tokens in each local window of the plurality of local window are able to attend to their corresponding global query tokens.) and the first image patches (Paragraphs [0051-52], Fig. 3, Hatamizadeh teaches the image patches are processed through the series of stages 304A-D of the vision transformer. Each stage 304 A-D includes alternating local self-attention and global self-attention modules to extract spatial features.).
In regards to Claim 10, Vaswani in view of Hatamizadeh teaches the image processing method of claim 1, wherein the image to be processed comprises a plurality of frames (Paragraph [0034], Vaswani teaches a batch of multiple input images and the layer input includes a respective index along the batch dimension for each input image in the batch. The Examiner interprets “multiple” images to be a plurality of frames since images and frames are synonymous.),
wherein the determining the recognition result of the image to be processed comprises determining a recognition result of the plurality of frames, respectively (Paragraphs [0018], [0034], Vaswani teaches the neural network can be configured to process an image to generate a classification output that includes a respective score corresponding to each of multiple categories. The score for a category indicates the likelihood that the image belongs to the category. The neural network may process a batch of multiple images (i.e., frames).), and
wherein the method further comprises:
based on the determined recognition result of the plurality of frames (Paragraphs [0018], [0034], Vaswani teaches the neural network can be configured to process an image to generate a classification output that includes a respective score corresponding to each of multiple categories. The score for a category indicates the likelihood that the image belongs to the category. The neural network may process a batch of multiple images (i.e., frames).), recognizing one or more highlight among the plurality of frames (Paragraphs [0018], [0034], Vaswani teaches the categories may represent global image properties (e.g., whether the image depicts a scene in the day or at night, or whether the image depicts a scene in the summer of the winter), and the image may belong to the category if it has the global property corresponding to the category. The neural network may process a batch of multiple images (i.e., frames). The Examiner interprets classifying images into specific categories to be recognizing a highlight since the category/classification can be of special interest i.e., predefined categories.).
In regards to Claim 11, Vaswani in view of Hatamizadeh teaches the image processing method of claim 10, wherein the recognizing the one or more highlights among the plurality of frames comprises:
dividing the image to be processed into a plurality of snippets of a fixed length (Paragraphs [0037-38], Fig. 2, Vaswani teaches blocking the block input “image” (210) into (H/b x W/b) non-overlapping (b, b, c) blocks, where b is a block size value for the layer. In Fig. 2, b is equal to 2 and the system has divided the input “image” (210) into a “blocked” image that includes four 2 by 2 element query blocks (220). The Examiner interprets the “image” (210) is processed into a plurality (e.g., 4) “blocks” of a fixed length (each is 2 blocks long). Further, under Broadest Reasonable Interpretation, the Examiner interprets “snippets” to mean a small piece, fragment, or brief extract of a larger work, text, or media. Therefore, a “block” or segment of an image is being interpreted as a “snippet” of the image.);
determining a highlight recognition score (Paragraph [0018], Vaswani teaches generating a respective score corresponding to each of multiple categories. The score for a category indicates a likelihood that the image belongs to the category.) for the plurality of snippets, respectively (Paragraph [0019], Vaswani teaches the neural network can be configured to process an image to generate a pixel-level classification output that includes, for each pixel, a respective score corresponding to each of multiple categories. For a given pixel, the score for a category indicates a likelihood that pixel belongs to the category. The Examiner interprets a “snippet” could be a singular pixel since it is a portion of the image and the claim does not specify that the “snippet” needs to be of a certain length.); and
classifying the plurality of snippets (Paragraph [0019], Vaswani teaches generating pixel-level classification output.) into highlight portions or non-highlight portions (Paragraph [0018], Vaswani teaches the categories may represent global image properties (e.g., whether the image depicts a scene in the day or at night).) based on the highlight recognition score (Paragraph [0019], Vaswani teaches the score for a category indicates a likelihood that pixel belongs to the category.).
In regards to Claim 13, Vaswani in view of Hatamizadeh teaches the image processing method of claim 3, wherein the at least one global token corresponds to one or more features to be extracted in the image to be processed (Paragraphs [0034-35], Hatamizadeh teaches global features may be key features detected within the input image. The global features may be extracted from the entirety of the input image by a global token generator. The global features may be used as a global query token.), and
wherein the generating the at least one kernel (Paragraphs [0050-51], Vaswani teaches the local self-attention layer can process layer input to generate a given context block by performing convolution using a kernel that includes a ‘1’ at each location corresponding to an element in the given context block.) for the image to be processed (Paragraphs [0033-34], Fig. 2, reference character 210, Vaswani teaches the layer input (210) is a [4, 4, c] “image”.), by the kernel generator (Paragraphs [0050-51], Vaswani teaches applying a 3D convolution that has a kernel made up of ones and zeros. The Examiner interprets a “kernel generator” to be a convolutional layer in light of Applicant’s specification, which states, “the kernel generator may employ, but is not limited to, convolutional layers, fully connecting layers, etc. It should be understood that, for some cases, it may be equivalent to use convolution layer or a fully connecting layer (Paragraph [0091]).”), comprises:
adapting a size of the at least one kernel to correspond to the one or more features to be extracted (Paragraphs [0050-51], Vaswani teaches the kernel can be the same size as the local window. In some implementations, the kernel has more elements than a local window, and each location that does not correspond to an element in the local window is a ‘0’. The kernel can be a sparse kernel or a one-hot kernel.).
In regards to Claim 14, Vaswani in view of Hatamizadeh teaches the image processing method of claim 1, wherein a resolution of the second image patches is lower than a resolution of the first image patches (Paragraph [0036], Hatamizadeh teaches for each local window and each of a plurality of sequential stages of the vision transformer, local and global self-attention may be computed for the input image. At each stage, or each of the plurality of stages, the vision transformer outputs feature representations for the input image. In an embodiment with a plurality of stages, a spatial resolution may be decreased after one or more of the stages of the vision transformer. The spatial resolution may be decreased by a downsampling block of the vision transformer.).
In regards to Claim 15, Vaswani discloses:
An image processing apparatus (Abstract, Vaswani teaches an apparatus for processing images using a computer vision neural network.), comprising:
at least one processor (Paragraph [0079], Vaswani teaches a central processing unit.); and
at least one memory storing instructions executable by the at least one processor (Paragraph [0079], Vaswani teaches a read only memory or a random access memory from which a central processing unit will receive instructions and data.),
wherein, by executing the instructions, the at least one processor is configured to control (Paragraph [0079], Vaswani teaches a central processing unit for performing or executing instructions.):
a first obtaining module (Paragraph [0034], Fig. 1, Vaswani teaches a neural network.) to obtain first image patches (Paragraphs [0033-34], Fig. 2, Vaswani teaches the neural network processes a batch of multiple input images in parallel and the layer input (210), which is a [4, 4, c] “image” includes a respective index along the batch dimension for each input image in the batch. The Examiner interprets the layer input (210), such as shown in Fig. 2 to be “first image patches”. Although Fig. 2 shows a single “image”, the neural network processes a batch of multiple input images in parallel to generate layer input for each input image in the batch, hence the plural “patches”.) corresponding to an image to be processed (Paragraph [0034], Vaswani teaches the layer input (210) generated from an input image.);
a first processing module (Paragraphs [0035-37], Vaswani teaches the local self-attention layer can group the elements of the corresponding layer input into multiple inputs called “query blocks.” The blocking is performed by the local self-attention layer.) to:
divide the first image patches into at least two groups via a window self-attention network (Paragraphs [0035-38], Vaswani teaches each local self-attention layer can group the elements of the corresponding layer input into multiple groups called “query blocks”. Block input (210) is divided into (H/b x W/b) non-overlapping (b, b, c) blocks where b is a block size value for the layer.),
obtain second image patches (Paragraph [0043], Fig. 2, reference character 250, Vaswani teaches a block attention output (250) for the given query block (220) that includes respective attention output for each element in the query block (220).) comprising local attention information (Paragraphs [0043], Fig. 2, Vaswani teaches the attention is local within the query block.); and
a first recognition module to determine a recognition result of the image to be processed (Paragraphs [0017-20], Vaswani teaches the neural network (150) can be configured to process an image to generate a pixel-level classification that includes, for each pixel, a respective score corresponding to each of multiple categories. The categories may be classes of objects, and a pixel may belong to a category if it is part of an object included in the object class corresponding to a category. The pixel-level classification output may be a semantic segmentation output. As another example, the neural network (150) can be configured to process an image to generate a regression output that estimated one or more continuous variables that characterize the image. The regression output may estimate coordinates of bounding boxes that enclose respective objects depicted in the image.), based on the second image patches (Paragraphs [0069], [0017], [0022], [0025], Vaswani teaches a block attention output (reference character 250 in Fig. 2), comprising determining a respective query for each element in the query block (220), determining a respective key for each element of the corresponding context block (240), determining a value for each element of the corresponding context block, and using the determined query, keys, and values to generate a respective attention output for each element of the query block. The feature representation of the input image can be one or more tensor numeric values that represent learned properties of the input image such as a single feature map. The output neural network maps the feature representation to an appropriate output for the computer vision task. The output for the computer vision task is one or more of a predicted classification, semantic segmentation, or an object detection for the input image.).
Vaswani does not explicitly disclose the following:
determine global attention information among first image patches in the at least two groups of the first image patches, respectively.
Hatamizadeh is in the same field of art of extracting local features from local windows/patches of an image. Further, Hatamizadeh teaches determine global attention information (Paragraphs [0015], [0059], Fig. 5B, Hatamizadeh teaches computing global attention for an image.) among first image patches in the at least two groups of the first image patches, respectively (Paragraphs [0060], [0059], [0070], [0033], Fig. 5B, Hatamizadeh teaches the image is split into a plurality of local windows. However, in order to facilitate long-range dependencies, Fig. 5B illustrates how global-attention is computed to allow cross-patch communication with those patches far beyond the local window. Global-self attention attends other regions (outside the local window) in the image. A global self-attention module that accesses, per local window of the plurality of local windows within the input image, global features extracted from an entirety of an input image, or from at least a portion of the input image outside the local window. The global features may be of any defined category (e.g., textures, shape descriptors, etc.).).
Therefore, it would have been obvious to one having ordinary skill in the art before the effective filing date of the claimed invention to modify the invention of Vaswani by computing global attention/extracting global features for each window within the input image that is taught by Hatamizadeh, to make the invention that computes both local and global attention information for the image patches of the image; thus, one of ordinary skilled in the art would be motivated to combine the references since there is a need for vision transformers to be able to capture long-range spatial dependencies in a less computationally expensive manner (Hatamizadeh, Paragraph [0005]). In addition, by computing both local and global self-attention during image processing, both short-range and long-range spatial dependencies may be respectively modeled by the vision transformer, which improves the quality of the feature representations obtained by the vision transformer.
Thus, the claimed subject matter would have been obvious to a person having ordinary skill in the art before the effective filing date of the claimed invention.
In regards to Claim 16, Vaswani in view of Hatamizadeh discloses the image processing apparatus of claim 15, wherein the at least one processor (Paragraph [0074], Vaswani teaches the data processing apparatus includes a programmable processor, or multiple processors.) is further configured to control:
the first processing module (Paragraph [0052], Hatamizadeh teaches a global self-attention module.) to:
determine, by a global token generator (Paragraph [0052], Fig. 3, Hatamizadeh teaches a global token generator (306), which is a CNN-like module that extracts features from the entire image only once every stage.), at least one global token corresponding to at least part of the image to be processed (Paragraphs [0052], [0061-64], Fig. 6, Hatamizadeh teaches the global token generator extracts features from the entire image only once every stage (304A-D). The global token generator generates global query tokens that encompass information across the entire input feature map for an input image.), based on the first image patches (Paragraph [0049-52], Hatamizadeh teaches the vision transformer (300) includes a stem layer to which an image is input. The stem layer obtains image patches for the image and projects those image patches into an embedding space having a defined dimension. The projected image patches are output from the stem layer and processed through a series of stages (304A-D) of the vision transformer (300). Each stage includes alternating local self-attention and global self-attention modules to extract spatial features.); and
the first recognition module (Paragraph [0045], Hatamizadeh teaches another processing block of the vision transformer for a computer vision task that is downstream from the vision transformer.) to:
determine the recognition result of the image to be processed (Paragraphs [0043], Hatamizadeh teaches the feature representations may be output to a downstream task, such as a computer vision-related downstream task. For example, the feature representations may be processed by the downstream task for performing image segmentation and/or object detection.), based on the at least one global token (Paragraphs [0034], [0052-53], Fig. 3, Hatamizadeh teaches the global features may be extracted from the entirety of the input image by a global token generator of the vision transformer. The global token generator extracts features from the entire image only once at every stage (304A-D). Resulting features output from the final stage 304D are passed through an average pooling layer and linear layer to create an embedding for a downstream task. The Examiner interprets since the global tokens are computed for each stage and contribute to the generation of the output embedding used for the downstream task, that the recognition result determined (via the downstream task) is “based on” the at least one global token.) and the second image patches (Paragraphs [0030-32], [0036-39], Fig. 3, Hatamizadeh teaches the input image is processed through each stage in the at least one stage. For each local window and each stage of the vision transformer, local and global self-attention may be computed for the input image. Each stage of the plurality of stages (see Fig. 3) of the vision transformer outputs feature representations for the input image. The Examiner interprets the output of the first stage to be the “second image patches”. Since the feature representations are output as embeddings for the input image and the feature representations may be output to a downstream task, such as classification, object detection, etc., the Examiner interprets the “recognition result” (result of the downstream task such as object detection) is “based on” the second image patches.).
In regards to Claim 17, Vaswani in view of Hatamizadeh discloses the image processing apparatus of claim 16, wherein the global token generator (Paragraphs [0061-64], Fig. 6, Hatamizadeh teaches a global token generator.) comprises a kernel generator (Paragraphs [0050-51], Vaswani teaches applying a 3D convolution that has a kernel made up of ones and zeros. The Examiner interprets a “kernel generator” to be a convolutional layer in light of Applicant’s specification, which states, “the kernel generator may employ, but is not limited to, convolutional layers, fully connecting layers, etc. It should be understood that, for some cases, it may be equivalent to use convolution layer or a fully connecting layer (Paragraph [0091]).”),
wherein the at least one processor (Paragraph [0074], Vaswani teaches a programmable processor, or multiple processors.) is further configured to control:
the first processing module (Paragraph [0050], Vaswani teaches the local self-attention layer can process the layer input to generate a given context block by performing convolution using a kernel. The Examiner interprets the local self-attention layer to be the first processing module.) to:
generate at least one kernel for the image to be processed (Paragraphs [0050-51], Vaswani teaches performing convolution using a kernel; that includes a ‘1’ at each location corresponding to an element in the given context block.), by the kernel generator (Paragraphs [0050-51], Vaswani teaches applying a 3D convolution that has a kernel made up of ones and zeros. The Examiner interprets a “kernel generator” to be a convolutional layer in light of Applicant’s specification, which states, “the kernel generator may employ, but is not limited to, convolutional layers, fully connecting layers, etc. It should be understood that, for some cases, it may be equivalent to use convolution layer or a fully connecting layer (Paragraph [0091]).”); and
the first recognition module (Paragraph [0045], Hatamizadeh teaches another processing block of the vision transformer for a computer vision task that is downstream from the vision transformer.) to:
determine the at least one global token (Paragraphs [0061- 63], Fig. 6, Hatamizadeh teaches generating global query tokens.) respectively corresponding to the at least one kernel (Paragraphs [0063], [0074] Hatamizadeh teaches global query tokens for interaction with local key and value features per local window when computing global self-attention.), based on the at least one kernel (Paragraphs [0063], [0074], Hatamizadeh teaches global query tokens that spatially correspond with local tokens of each local window of a plurality of local windows within the image, such that the local tokens in each local window of the plurality of local window are able to attend to their corresponding global query tokens.) and the first image patches (Paragraphs [0051-52], Fig. 3, Hatamizadeh teaches the image patches are processed through the series of stages 304A-D of the vision transformer. Each stage 304 A-D includes alternating local self-attention and global self-attention modules to extract spatial features.).
In regards to Claim 20, Vaswani discloses the following:
A non-transitory computer readable storage medium (Claim 11, Vaswami teaches one or more non-transitory computer-readable storage media.) having a computer program stored therein, wherein when the computer program is executed by at least one processor, the computer program performs executed by at least one processor, the computer program performs the image processing method (Claim 11, Paragraph [0079], Vaswami teaches storing instructions that when executed by one or more computers cause the one or more computers to implement a neural network having one or more local self-attention layers. Computers suitable for the execution of the computer program can be based on general or special purpose microprocessors or both, or any other kind of central processing unit.)
Vaswani does not explicitly disclose according to claim 1. Examiner states see method of combining references for Claim 1.
Claims 4-6 and 18-19 are rejected under 35 U.S.C. 103(a) as being unpatentable over Vaswani et al. (U.S. Patent Pub. No. 20210390410 A1, hereafter referred to as Vaswani) in view of Hatamizadeh et al. (U.S. Patent Pub No. 20230394781 A1, hereafter referred to as Hatamizadeh) in further view of Zhang et al. (U.S. Patent Pub. No. 20230306600 A1, hereafter referred to as Zhang).
Regarding Claim 4, Vaswani in view of Hatamizadeh discloses the image processing method of claim 2, wherein the determining the recognition result of the image to be processed based on the at least one global token and the second image patches, comprises:
determining, (Paragraph [0060], Fig. 5B, Hatamizadeh teaches global self-attention is computed to allow cross-patch communication with those patches far beyond the local window. Global self-attention attends other regions (outside the local window) in the image via a global query token that represents an image embedding extracted with CNN-like module. The global features are extracted from the entire input features, and then are repeated to form global query tokens. The global query token is interacted with local key and value tokens (per local window), hence allowing the capture of long-range information via cross-region interaction.);
PNG
media_image7.png
207
495
media_image7.png
Greyscale
obtaining third image patches comprising the global attention information and the local attention information (Paragraphs [0049-53], Fig. 3, Hatamizadeh teaches a multi-stage vision transformer (300) configured to provide global context and downsampling. The projected image patches are output from the stem layer (302) and processed through a series of stages (304A-D) of a vision transformer. Each stage (304A-D) includes alternating local self-attention and global self-attention modules to extract spatial features. Feature representations are obtained at several resolutions (one per stage 304A-D). See annotated Fig. 3 for the Examiner’s interpretation of “third image patches”, specifically the output of the down sampling stage after the first stage of the transformer (304A).);
PNG
media_image8.png
522
1067
media_image8.png
Greyscale
and
determining the recognition result of the image to be processed (Paragraphs [0024], [0039], [0043], Hatamizadeh teaches the downstream task can then process the given input to provide, for example, image classification, object detection, instance segmentation, semantic segmentation, or other computer vision-related information for the input image.) based on the third image patches (Paragraph [0053], Fig. 3, Hatamizadeh teaches resulting features output from the final stage (304D) are passed through an average pooling layer (310) and then a linear layer (312) to create an embedding for a downstream task. The Examiner interprets the third image patches (see annotated Fig. 3) are inputted into the subsequent stage (304B) of the vision transformer, which extracts features and feeds them to the following stage (304C) and the final stage (304D) which outputs the resulting features output which is used for the downstream task (object detection, etc.). Therefore, the Examiner interprets the recognition result (such as object detection in an image) is “based on” the third image patches.).
Vaswani in view of Hatamizadeh does not explicitly disclose determining, via a cross-attention network, (attention information).
[AltContent: arrow]Zhang is in the same field of art of generating attention information to perform a downstream task such as semantic segmentation, for example. Further, Zhang discloses determining, via a cross-attention network, attention information (Paragraph [0021], Fig. 8, Zhang teaches a transformer-based cross-attention neural network. See Fig. 8 below, cross attention source features.).
PNG
media_image9.png
729
1122
media_image9.png
Greyscale
Therefore, it would have been obvious to one having ordinary skill in the art before the effective filing date of the claimed invention to modify the invention of Vaswani in view of Hatamizadeh by using a cross-attention network to perform an image processing recognition task such as semantic image segmentation that is taught by Zhang, to make the invention that performs semantic image segmentation to recognize/identify objects/things in video frames/images using a machine learning system, specifically one that includes one or more cross-attention transformer layers; thus, one of ordinary skilled in the art would be motivated to combine the references since the cross-attention layers can be provided with more accurate and comprehensive information that ultimately drives the generation of semantic segmentation masks with improved accuracy and efficiency (Zhang, Paragraph [0092]). In addition, by aligning queries from one sequence with relevant elements in another, cross-attention ensures that the model focuses on the most critical information.
Thus, the claimed subject matter would have been obvious to a person having ordinary skill in the art before the effective filing date of the claimed invention.
In regards to Claim 5, Vaswani in view of Hatamizadeh in further view of Zhang discloses the image processing method of claim 4, wherein at least one of the window self-attention network, the global token generator and the cross-attention network comprises a first neural network (Paragraph [0004], Vaswani teaches the computer vision neural network includes one or more local self-attention vision neural network layers. Under Broadest Reasonable Interpretation, the Examiner interprets “at least one of” to mean only one of the three: window self-attention network, the global token generator or the cross-attention network comprises a first neural network is required to meet the claim limitation.),
wherein the determining the recognition result of the image to be processed based on the first image patches comprises:
obtaining fourth image patches comprising the global attention information and the local attention information based on the first image patches (Paragraphs [0051-51], Fig. 3, Hatamizadeh teaches the projected image patches are output from the stem layer (Fig. 3, reference character 302) and processed through a series of stages of the vision transformer. Each stage (304A-D) includes alternating local self-attention and global self-attention modules to extract spatial features. Both local self-attention and global self-attention modules operate in local windows of the image. See annotated Fig. 3 below. The Examiner is interpreting the output patches of the second stage (304B) as “fourth image patches”. Further, the Examiner asserts that the fourth image patches comprise global attention and local attention information since each stage includes alternating local self-attention and global self-attention modules to extract spatial features.), via at least one first neural network (Paragraph [0024], Fig. 3, Hatamizadeh teaches a vision transformer (e.g., neural network, deep learning model) is configured to process images, using both local and global self-attention, to derive information from those images.);
PNG
media_image8.png
522
1067
media_image8.png
Greyscale
and
determining the recognition result of the image to be processed (Paragraphs [0024], [0039], [0043], Hatamizadeh teaches the downstream task can then process the given input to provide, for example, image classification, object detection, instance segmentation, semantic segmentation, or other computer vision-related information for the input image.), based on the fourth image patches (Paragraph [0053], Fig. 3, Hatamizadeh teaches resulting features output from the final stage (304D) are passed through an average pooling layer (310) and then a linear layer (312) to create an embedding for a downstream task. The Examiner interprets the fourth image patch (see annotated Fig. 3 above) is inputted into the subsequent stage (304C) of the vision transformer, which then extracts features and feeds them to the final stage (304D) which outputs the resulting features output which is used for the downstream task (object detection, etc.). Therefore, the Examiner interprets the recognition result (such as object detection in an image) is “based on” the fourth image patches.),
wherein, the obtaining the fourth image patches comprising the global attention information and the local attention information based on the first image patches, via the at least one first neural network, further comprises:
performing at least one down sampling for the first image patches (Paragraphs [0052-57], Fig. 3, Vaswani teaches an attention downsampling layer that reduces the spatial dimension of layer input (310), i.e., that “downsample” the layer input. Fig. 3 is an illustration of a downsampling local self-attention mechanism applied by an attention downsampling layer.),
PNG
media_image10.png
294
762
media_image10.png
Greyscale
and wherein the at least one down sampling respectively comprises:
down sampling output image patches (Paragraph [0056], Vaswani teaches the attention downsampling layer generates a respective query block (320) corresponding to each of the H/b×W/b non-overlapping (b, b, c) blocks. In the example of Fig. 3, the downsampling factor is 2.) of a previous first neural network (Paragraph [0052], Hatamizadeh teaches following each stage (304 A-C), with the exception of the final stage (304D), is a downsampling block (308A-C). The downsampling block decreases a spatial resolution of the output of the immediate prior stage (304A-C) by 2 while increasing the number of channels.) to obtain down sampled results (Paragraphs [0052-56], Vaswani teaches the neural network includes a downsampling layer that reduces the spatial dimension of layer input.); and
inputting the down sampled results into a first neural network (Paragraphs [0022], [0025], Fig. 1, Vaswani teaches an output neural network (140) that processes the feature representation (130) to generate the output (152) for the computer vision task.) next to the previous first neural network (Paragraphs [0015], [0022], [0025], Fig. 1, Vaswani teaches a computer vision neural network (150) includes a backbone neural network (110) that processes the input image (102) to generate a feature representation (130) of the input image (102) and an output neural network (140) that processes the feature representation (130) to generate output (152) for the computer vision task. The feature representation (130) can be a single feature map having smaller spatial dimensions than the input image (i.e., down sampled) but with a larger number of channels than the input image. The Examiner interprets the output neural network (140) and the backbone neural network (110) are “next to” each other since there are now additional components depicted in Fig. 1 between them and the feature representation (130) from the backbone neural network (110) are processed by the output neural network (140).).
PNG
media_image11.png
628
563
media_image11.png
Greyscale
In regards to Claim 6, Vaswani in view of Hatamizadeh in further view of Zhang discloses the image processing method of claim 5, wherein, for each of the output image patches of the previous first neural network (Paragraph [0046], Figs. 2 & 3, Hatamizadeh teaches the first stage processes the first input to generate a first output, and the first output is provided as second input to the second stage (202B) of the vision transformer for processing. Thus, while the first stage (202A) processes the image, each of the subsequent stages 202A-N of the vision transformer (200) process the output of the immediate prior stages (202A-N).), the down sampling the output image patches comprises (Paragraphs [0048], [0052-53], Fig. 3, Hataizadeh teaches the vision transformer (200) may include additional processing blocks situated between one or more of the plurality of stages (202A-N), such as downsampling blocks. For example, the downsampling block (308A-C) decreases a spatial resolution of the output of the immediate prior stage (304A-C) by 2 while increasing a number of channels.):
grouping feature points of each of the output image patches into grouped feature maps (Paragraphs [0053], [0055], [0062], [0075], Hatamizadeh teaches feature representations are obtained at several resolutions (one per stage 304A-D) by decreasing the spatial dimensions while expanding the embedding dimension (e.g., by factors of 2 and 2, respectively).); and
concatenating the grouped feature maps in a channel dimension to obtain connected feature maps (Paragraph [0066], Hatamizadeh teaches multi-head attention is employed and the outputs are concatenated and projected into the expected dimension.).
In regards to Claim 18, Vaswani in view of Hatamizadeh discloses the image processing apparatus of claim 16, wherein the at least one processor (Paragraph [0074], Vaswani teaches the data processing apparatus includes a programmable processor, or multiple processors.) is further configured to control:
the first processing module (Paragraph [0052], Hatamizadeh teaches a global self-attention module.) to:
determine, via a (Paragraph [0060], Fig. 5B, Hatamizadeh teaches global self-attention is computed to allow cross-patch communication with those patches far beyond the local window. Global self-attention attends other regions (outside the local window) in the image via a global query token that represents an image embedding extracted with CNN-like module. The global features are extracted from the entire input features, and then are repeated to form global query tokens. The global query token is interacted with local key and value tokens (per local window), hence allowing the capture of long-range information via cross-region interaction.), and
obtain third image patches comprising the global attention information and the local attention information (Paragraphs [0049-53], Fig. 3, Hatamizadeh teaches a multi-stage vision transformer (300) configured to provide global context and downsampling. The projected image patches are output from the stem layer (302) and processed through a series of stages (304A-D) of a vision transformer. Each stage (304A-D) includes alternating local self-attention and global self-attention modules to extract spatial features. Feature representations are obtained at several resolutions (one per stage 304A-D). See annotated Fig. 3 for the Examiner’s interpretation of “third image patches”, specifically the output of the down sampling stage after the first stage of the transformer (304A).); and
the first recognition module (Paragraph [0045], Hatamizadeh teaches another processing block of the vision transformer for a computer vision task that is downstream from the vision transformer.) to:
determine the recognition result of the image to be processed (Paragraphs [0024], [0039], [0043], Hatamizadeh teaches the downstream task can then process the given input to provide, for example, image classification, object detection, instance segmentation, semantic segmentation, or other computer vision-related information for the input image.) based on the third image patches (Paragraph [0053], Fig. 3, Hatamizadeh teaches resulting features output from the final stage (304D) are passed through an average pooling layer (310) and then a linear layer (312) to create an embedding for a downstream task. The Examiner interprets the third image patches (see annotated Fig. 3) are inputted into the subsequent stage (304B) of the vision transformer, which extracts features and feeds them to the following stage (304C) and the final stage (304D) which outputs the resulting features output which is used for the downstream task (object detection, etc.). Therefore, the Examiner interprets the recognition result (such as object detection in an image) is “based on” the third image patches.).
Vaswani in view of Hatamizadeh does not explicitly disclose determining, via a cross-attention network, (attention information).
Zhang is in the same field of art of generating attention information to perform a downstream task such as semantic segmentation, for example. Further, Zhang discloses determining, via a cross-attention network, attention information (Paragraph [0021], Fig. 8, Zhang teaches a transformer-based cross-attention neural network.).
Therefore, it would have been obvious to one having ordinary skill in the art before the effective filing date of the claimed invention to modify the invention of Vaswani in view of Hatamizadeh by using a cross-attention network to perform an image processing recognition task such as semantic image segmentation that is taught by Zhang, to make the invention that performs semantic image segmentation to identify a recognition result i.e., identify an object, etc. using a machine learning system, specifically one that includes one or more cross-attention transformer layers; thus, one of ordinary skilled in the art would be motivated to combine the references since the cross-attention layers can be provided with more accurate and comprehensive information that ultimately drives the generation of semantic segmentation masks with improved accuracy and efficiency (Zhang, Paragraph [0092]). In addition, by aligning queries from one sequence with relevant elements in another, cross-attention ensures that the model focuses on the most critical information.
Thus, the claimed subject matter would have been obvious to a person having ordinary skill in the art before the effective filing date of the claimed invention.
In regards to Claim 19, Vaswani in view of Hatamizadeh in further view of Zhang discloses the image processing apparatus of claim 18, wherein at least one of the window self-attention network, the global token generator and the cross-attention network comprises a first neural network (Paragraph [0004], Vaswani teaches the computer vision neural network includes one or more local self-attention vision neural network layers. Under Broadest Reasonable Interpretation, the Examiner interprets “at least one of” to mean only one of the three: window self-attention network, the global token generator or the cross-attention network comprises a first neural network is required to meet the claim limitation.),
wherein the at least one processor (Paragraph [0074], Vaswani teaches a programmable processor or multiple processors.) is further configured to:
for determining the recognition result of the image to be processed (Paragraphs [0017-20], Vaswani teaches the system can process the input image (102) to generate a corresponding output, e.eg., a classification output, a regression output, or a combination thereof.) based on the first image patches (Paragraphs [0017], [0033-34], Fig. 2, Vaswani teaches processing an input that includes an image to generate a corresponding output, e.g., classification output, etc. The Examiner interprets input image (102) to be a “first image patch”. Further, the Examiner asserts that there is multiple “first image patches” since the neural network processes a batch of multiple input images in parallel.), control the first processing module to:
obtain fourth image patches comprising the global attention information and local attention information based on the first image patches (Paragraphs [0051-51], Fig. 3, Hatamizadeh teaches the projected image patches are output from the stem layer (Fig. 3, reference character 302) and processed through a series of stages of the vision transformer. Each stage (304A-D) includes alternating local self-attention and global self-attention modules to extract spatial features. Both local self-attention and global self-attention modules operate in local windows of the image. See annotated Fig. 3 below. The Examiner is interpreting the output patches of the second stage (304B) as “fourth image patches”. Further, the Examiner asserts that the fourth image patches comprise global attention and local attention information since each stage includes alternating local self-attention and global self-attention modules to extract spatial features.), via at least one first neural network (Paragraph [0024], Fig. 3, Hatamizadeh teaches a vision transformer (e.g., neural network, deep learning model) is configured to process images, using both local and global self-attention, to derive information from those images.);
for determining the recognition result of the image to be processed (Paragraphs [0024], [0039], [0043], Hatamizadeh teaches the downstream task can then process the given input to provide, for example, image classification, object detection, instance segmentation, semantic segmentation, or other computer vision-related information for the input image.) based on the first image patches (Paragraph [0053], Fig. 3, Hatamizadeh teaches resulting features output from the final stage (304D) are passed through an average pooling layer (310) and then a linear layer (312) to create an embedding for a downstream task. The Examiner interprets the fourth image patch (see annotated Fig. 3 above) is inputted into the subsequent stage (304C) of the vision transformer, which then extracts features and feeds them to the final stage (304D) which outputs the resulting features output which is used for the downstream task (object detection, etc.). Therefore, the Examiner interprets the recognition result (such as object detection in an image) is “based on” the fourth image patches.), control the first recognition module (Paragraphs [0017-20], Vaswani teaches the neural network (150) can be configured to process an image to generate a classification.) to:
determine the recognition result of the image to be processed (Paragraphs [0024], [0039], [0043], Hatamizadeh teaches the downstream task can then process the given input to provide, for example, image classification, object detection, instance segmentation, semantic segmentation, or other computer vision-related information for the input image.), based on the fourth image patches (Paragraph [0053], Fig. 3, Hatamizadeh teaches resulting features output from the final stage (304D) are passed through an average pooling layer (310) and then a linear layer (312) to create an embedding for a downstream task. The Examiner interprets the fourth image patch (see annotated Fig. 3 above) and is inputted into the subsequent stage (304C) of the vision transformer, which then extracts features and feeds them to the final stage (304D) which outputs the resulting features output which is used for the downstream task (object detection, etc.). Therefore, the Examiner interprets the recognition result (such as object detection in an image) is “based on” the fourth image patches.);
for obtaining the fourth image patches (Paragraph [0051], Fig. 3, Hatamizadeh teaches the projected image patches are output from the stem layer and processed through a series of stages (304A-D) of the vision transformer. See annotated Fig. 3 below.)
PNG
media_image12.png
522
1171
media_image12.png
Greyscale
comprising the global attention information and the local attention information (Paragraphs [0036], [0047], Hatamizadeh teaches for each local window and each stage of the vision transformer, local and global self-attention may be computed for the input image.) based on the first image patches (Paragraphs [0051], [0045-46], Fig. 3, Hatamizadeh teaches the projected image patches are processed through a series of stages. The processing stages (304A-D) operate sequentially (i.e., the first stage processes the first input to generate a first output, and the first output is provided as input to the second stage of the vision transformer for processing). Therefore, the Examiner interprets the fourth image patches are “based” on the first image patches output by the stem layer since they have been processed through a series of stages.) via the at least one first neural network (Paragraph [0024], Hatamizadeh teaches a vision transformer (e.g., neural network, deep learning model).), control the first processing module (Paragraph [0052], Vaswani teaches the neural network includes an attention downsampling layer that reduce the spatial dimension of layer input, i.e., that “downsample” the layer input.) to:
perform at least one down sampling for the first image patches (Paragraphs [0052-57], Fig. 3, Vaswani teaches an attention downsampling layer that reduces the spatial dimension of layer input (310), i.e., that “downsample” the layer input. Fig. 3 is an illustration of a downsampling local self-attention mechanism applied by an attention downsampling layer.); and for performing the at least one down sampling (Paragraphs [0052-57], Vaswani teaches downsampling the layer input.), control the first processing module (Paragraph [0052], Vaswani teaches the neural network includes an attention downsampling layer that reduce the spatial dimension of layer input, i.e., that “downsample” the layer input.) to:
down sample output image patches (Paragraph [0056], Vaswani teaches the attention downsampling layer generates a respective query block (320) corresponding to each of the H/b×W/b non-overlapping (b, b, c) blocks. In the example of Fig. 3, the downsampling factor is 2.) of a previous first neural network (Paragraph [0052], Hatamizadeh teaches following each stage (304 A-C), with the exception of the final stage (304D), is a downsampling block (308A-C). The downsampling block decreases a spatial resolution of the output of the immediate prior stage (304A-C) by 2 while increasing the number of channels.) to obtain down sampled results (Paragraphs [0052-56], Vaswani teaches the neural network includes a downsampling layer that reduces the spatial dimension of layer input.), and
input the down sampled results into a first neural network (Paragraphs [0022], [0025], Fig. 1, Vaswani teaches an output neural network (140) that processes the feature representation (130) to generate the output (152) for the computer vision task.) next to the previous first neural network(Paragraphs [0015], [0022], [0025], Fig. 1, Vaswani teaches a computer vision neural network (150) includes a backbone neural network (110) that processes the input image (102) to generate a feature representation (130) of the input image (102) and an output neural network (140) that processes the feature representation (130) to generate output (152) for the computer vision task. The feature representation (130) can be a single feature map having smaller spatial dimensions than the input image (i.e., down sampled) but with a larger number of channels than the input image. The Examiner interprets the output neural network (140) and the backbone neural network (110) are “next to” each other since there are now additional components depicted in Fig. 1 between them and the feature representation (130) from the backbone neural network (110) are processed by the output neural network (140).).
Claim 7 is rejected under 35 U.S.C. 103(a) as being unpatentable over Vaswani et al. (U.S. Patent Pub. No. 20210390410 A1, hereafter referred to as Vaswani) in view of Hatamizadeh et al. (U.S. Patent Pub No. 20230394781 A1, hereafter referred to as Hatamizadeh) in further view of Zhang et al. (U.S. Patent Pub. No. 20230306600 A1, hereafter referred to as Zhang) in further view of Bertasius et al. (U.S. Patent Pub. No. 20220253633 A1, hereafter referred to as Bertasius).
Regarding Claim 7, Vaswani in view of Hatamizadeh in further view of Zhang discloses the image processing method of claim 4, wherein the determining the recognition result of the image to be processed based on the third image patches comprises:
determining fifth image patches comprising the global attention information, the local attention information (Paragraphs [0036], [0046], Fig. 3, Hatamizadeh teaches for each local window and each stage of the vision transformer, local and global self-attention may be computed for the input image. For each local window and each of a plurality of (e.g., sequential) stages of the vision transformer, local and global self-attention may be computed for the input image. Each stage of the plurality of stages of the vision transformer outputs feature representations for the input image. See annotated Fig. 3 below for the Examiner’s interpretation of “fifth image patches”. The Examiner interprets the “fifth image patches” to be the downsampled patches output by downsampling block 308B in Fig. 3.)
PNG
media_image13.png
522
1135
media_image13.png
Greyscale
(Paragraph [0051], Fig. 3, Hatamizadeh teaches a series of stages (304A-D) of the vision transformer. Each stage includes alternating local self-attention and global self-attention modules to extract spatial features. The local self-attention module is composed of a local multi-head self-attention layer as well as corresponding multilayer perceptron. The global self-attention module is composed of a global MSA and corresponding MLP. Therefore, the Examiner interprets since each stage of the vision transformer (300) contains 2 modules, each of which contains an MSA layer and a MLP, the Examiner interprets the vision transformer architecture is made up of more than two neural networks (i.e., each stage includes MLP, a type of neural network).); and
determining the recognition result of the image to be processed (Paragraphs [0024], [0039], Hatamizadeh teaches the downstream task can then process the given input to provide, for example, image classification, object detection, instance segmentation, semantic segmentation, or other computer vision-related information for the input image.), based on the fifth image patches (Paragraphs [0045-46], [0053], Figs. 2 & 3, Hatamizadeh teaches the final output of the stages 202A-N includes the feature representations of the input image, which may in turn be provided to another processing block of the vision transformer (200) or a computer vision task that is downstream from the vision transformer (200). Each of the subsequent stages 202A-N of the vision transformer (200) processes the output of the immediate prior one of the stages 202A-N. The resulting features from the final stage are passed through an average pooling layer (310) and then a linear layer (312) to create an embedding for a downstream task.).
Vaswani in view of Hatamizadeh in further view of Zhang does not explicitly disclose (determining) temporal information based on the third image patches.
Bertasius is in the same field of art of capturing both local as well as global long-range dependencies for images/frames by directly comparing feature activations at all space-time locations. Further, Bertasius teaches (determining) temporal information based on the third image patches (Paragraph [0019], Fig. 3B, Bertasius teaches the patches (303) indicated by vertical stripes are used for computing either a spatial attention or a temporal attention for the patch (301). In Fig. 3B, a spatiotemporal attention may be calculated based on comparisons between embedding vector corresponding to the patch (301) generated at encoding block I-1 and embedding vectors corresponding to the other patches (303) in the frame t. The Examiner interprets since the temporal attention is computed using the striped patches, in which 15 striped patches are shown, that temporal information is based on “third image patches” (in addition to the other striped patches). The claim is silent to whether other image patches are also used as long as “third image patches” are used for determining temporal information.).
PNG
media_image14.png
222
362
media_image14.png
Greyscale
Therefore, it would have been obvious to one having ordinary skill in the art before the effective filing date of the claimed invention to modify the invention of Vaswani in view of Hatamizadeh in further view of Zhang by computing temporal attention for the patch based on comparisons with several patches within the same frame that is taught by Bertasius, to make the invention that generates temporal information across frames of a video; thus, one of ordinary skilled in the art would be motivated to combine the references since the space (spatial) attention only scheme neglects to capture temporal dependencies across frames. This approach leads to degraded classification accuracy compared to full spatiotemporal attention (Bertasius, Paragraph [0019]).
Thus, the claimed subject matter would have been obvious to a person having ordinary skill in the art before the effective filing date of the claimed invention.
Claim 12 is rejected under 35 U.S.C. 103(a) as being unpatentable over Vaswani et al. (U.S. Patent Pub. No. 20210390410 A1, hereafter referred to as Vaswani) in view of Hatamizadeh et al. (U.S. Patent Pub No. 20230394781 A1, hereafter referred to as Hatamizadeh) in further view of Suri et al. (U.S. Patent No. 10,664,687 B2, hereafter referred to as Suri).
Regarding Claim 12, Vaswani in view of Hatamizadeh discloses the image processing method of claim 11.
Vaswani in view of Hatamizadeh does not explicitly disclose wherein the recognizing the one or more highlights among the plurality of frames further comprises:
integrating adjacent snippets classified as a highlight portion, based on the adjacent snippets corresponding to a same type of highlight; and
determining a start time and an end time of the integrated adjacent snippets as the one or more highlights.
Suri is in the same field of art of identifying important video sections in a video file (i.e. “highlights”). Further, Suri teaches wherein the recognizing the one or more highlights among the plurality of frames further comprises:
integrating adjacent snippets (Col. 11, lines 5-50, Fig. 3, Suri teaches once the feature points are identified for two adjacent frames, the motion analysis module may determine a transform that aligns the two adjacent frames such that a maximum number of feature points match.) classified as a highlight portion (Col. 6, lines 7-11, Suri teaches feature points are interest points in a video frame that can be reliably located across multiple video frames.), based on the adjacent snippets corresponding to a same type of highlight (Col. 10, lines 55-67 through Col. 11, lines 1-4, Suri teaches locating feature points for two adjacent frames. To detect the feature points, the motion analysis module may downsample the image and create a pyramid of down sampled images of smaller dimensions. The down sampled images are compared by the motion analysis module to determine the common points (i.e., feature points) among the down sampled images.); and
determining a start time and an end time of the integrated adjacent snippets (Col. 14, lines 43-49, Suri teaches storing metadata regarding the video sections in the ranked video file. The metadata may include the start and end time of each video section.) as the one or more highlights (Col. 14, lines 38-45, Suri teaches the video ranking module may rank video sections of the video file based on their section importance values.).
Therefore, it would have been obvious to one having ordinary skill in the art before the effective filing date of the claimed invention to modify the invention of Vaswani in view of Hatamizadeh by integrating/aligning video frames corresponding to the same feature points in a video segment and determining start/end times for each video section that is taught by Suri, to make the invention that identifies and ranks the importance of video sections in a video file; thus, one of ordinary skilled in the art would be motivated to combine the references since presenting a user with importance information for sections of a video file enables the user to tell at a glance the interesting portions of the video file. This information may assist the user in editing the video file to improve content quality or highlight particular sections of the video file (Suri, Col. 21, lines 8-20).
Thus, the claimed subject matter would have been obvious to a person having ordinary skill in the art before the effective filing date of the claimed invention.
Pertinent Prior Art
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure.
Chen et al. (U.S. Patent Pub. No. 20230386197 A1) teaches a Vision Transformer (ViT) architecture for vision tasks, such as image classification, object detection, and action recognition applied to visual content such as images or video content. The ViT architecture generates regional tokens and local tokens from an image with different patch sizes, where each regional token is associated with a set of local tokens based on spatial location.
Arnab et al. (U.S. Patent No. 12,112,538 B2) teaches a computer-implemented method for classifying video data with improved accuracy includes obtaining video data comprising a plurality of video frames, extracting a plurality of video tokens from the video data, the plurality of tokens comprising a representation of spatiotemporal information in the video data, providing, by the computing system, the plurality of video tokens as input to a video understanding model, the video understanding model comprising video transformer encoder model; and receiving a classification output from the video understanding model.
Xu et al. (NPL “Long Short-Term Transformer for Online Action Detection”, 2021) teaches a “Long Short-term Transformer (LSTR)”, a temporal modeling algorithm for online action detection, which employs long- and short-term memory mechanism to model prolonged sequence data. The mechanism consists of an LSTR encoder that leverages coarse-scale historical information from an extended temporal window (e.g., 2048 frames spanning up to 8 minutes), together with an LSTR decoder that focuses on a short window (e.g., 32 frames spanning 8 seconds) to model the fine-scale characteristics of the data.
Allowable Subject Matter
Claims 8-9 are objected to as being dependent upon a rejected base claim, but would be allowable if rewritten in independent form including all of the limitations of the base claim and any intervening claims.
The following is a statement of reasons for the indication of allowable subject matter:
Regarding Claim 8 and dependents, no prior art teaches wherein the determining the fifth image patches comprising the global attention information, the local attention information and the temporal information based on the third image patches, via the second neural network, comprises: obtaining, from a predetermined short memory pool, sixth image patches corresponding to at least one frame of a processed image prior to the image to be processed; determining temporal third image patches comprising the global attention information, the local attention information and the temporal information, based on the third image patches and the sixth image patches; down sampling the third image patches to obtain seventh image patches; obtaining, from a predetermined long memory pool, eighth image patches corresponding to the at least one frame of the processed image prior to the image to be processed; determining temporal seventh image patches comprising temporal information, based on the seventh image patches and the eighth image patches; and obtaining the fifth image patches comprising the global attention information, the local attention information and the temporal information, based on the temporal third image patches and the temporal seventh image patches.
Conclusion
Any inquiry concerning this communication or earlier communications from the examiner should be directed to SYDNEY L BLACKSTEN whose telephone number is (571)272-7651. The examiner can normally be reached 8:30am-4:30pm.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Oneal Mistry can be reached at 313-446-4912. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/SYDNEY L BLACKSTEN/Examiner, Art Unit 2674
/ONEAL R MISTRY/Supervisory Patent Examiner, Art Unit 2674