DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Status of Claims
This communication is in response to the Application Filed on 10/11/2024
Claims 1–10 are pending in this application.
Drawings
The drawing(s) filed on 10/11/2024 are accepted by the Examiner.
Claim Interpretation
The following is a quotation of 35 U.S.C. 112(f):
(f) Element in Claim for a Combination. – An element in a claim for a combination may be expressed as a means or step for performing a specified function without the recital of structure, material, or acts in support thereof, and such claim shall be construed to cover the corresponding structure, material, or acts described in the specification and equivalents thereof.
The following is a quotation of pre-AIA 35 U.S.C. 112, sixth paragraph:
An element in a claim for a combination may be expressed as a means or step for performing a specified function without the recital of structure, material, or acts in support thereof, and such claim shall be construed to cover the corresponding structure, material, or acts described in the specification and equivalents thereof.
The claims in this application are given their broadest reasonable interpretation using the plain meaning of the claim language in light of the specification as it would be understood by one of ordinary skill in the art. The broadest reasonable interpretation of a claim element (also commonly referred to as a claim limitation) is limited by the description in the specification when 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is invoked.
As explained in MPEP § 2181, subsection I, claim limitations that meet the following three-prong test will be interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph:
(A) the claim limitation uses the term “means” or “step” or a term used as a substitute for “means” that is a generic placeholder (also called a nonce term or a non-structural term having no specific structural meaning) for performing the claimed function;
(B) the term “means” or “step” or the generic placeholder is modified by functional language, typically, but not always linked by the transition word “for” (e.g., “means for”) or another linking word or phrase, such as “configured to” or “so that”; and
(C) the term “means” or “step” or the generic placeholder is not modified by sufficient structure, material, or acts for performing the claimed function.
Use of the word “means” (or “step”) in a claim with functional language creates a rebuttable presumption that the claim limitation is to be treated in accordance with 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph. The presumption that the claim limitation is interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is rebutted when the claim limitation recites sufficient structure, material, or acts to entirely perform the recited function.
Absence of the word “means” (or “step”) in a claim creates a rebuttable presumption that the claim limitation is not to be treated in accordance with 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph. The presumption that the claim limitation is not interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is rebutted when the claim limitation recites function without reciting sufficient structure, material or acts to entirely perform the recited function.
Claim limitations in this application that use the word “means” (or “step”) are being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, except as otherwise indicated in an Office action. Conversely, claim limitations in this application that do not use the word “means” (or “step”) are not being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, except as otherwise indicated in an Office action.
This application includes one or more claim limitations that do not use the word “means,” but are nonetheless being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, because the claim limitation(s) uses a generic placeholder that is coupled with functional language without reciting sufficient structure to perform the recited function and the generic placeholder is not preceded by a structural modifier.
Such claim limitation(s) is:
“a construction module” in claim(s) 7
“a prediction module” in claim(s) 7
“a training module” in claim(s) 7
Because this claim limitation(s) is being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, it is being interpreted to cover the corresponding structure described in the specification as performing the claimed function, and equivalents thereof.
Claim(s) 7: “a construction module” corresponds to FIG. 6 – element 610. “a construction module 610, which is used for constructing a training dataset, wherein the training dataset comprises: a prediction result and an annotation result of a rotated bounding box in the prediction result, the annotation result comprising pixel coordinates and a rotation angle of the rotated bounding box”, Applicant Specification ¶ [0079].
Claim(s) 7: “a prediction module” corresponds to FIG. 6 – element 620. “a prediction module 620, which is used for inputting the prediction result into a rotated bounding box object detection model to be trained, and obtaining a prediction result, wherein the prediction result comprises: a predicted pixel coordinate value and a predicted angle value”, Applicant Specification ¶ [0079].
Claim(s) 7: “a training module” corresponds to FIG. 6 – element 630. “a training module 630, which is used for comparing the prediction result with the annotation result, using a loss function to optimize the rotated bounding box object detection model to be trained, and obtaining a trained rotated bounding box object detection model.”, Applicant Specification ¶ [0079].
If applicant does not intend to have this limitation(s) interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, applicant may: (1) amend the claim limitation(s) to avoid it being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph (e.g., by reciting sufficient structure to perform the claimed function); or (2) present a sufficient showing that the claim limitation(s) recite(s) sufficient structure to perform the claimed function so as to avoid it being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph.
Claim Rejections - 35 USC § 101
35 U.S.C. 101 reads as follows:
Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title.
The text of those sections of Title 35, U.S. Code not included in this action can be found in a prior Office action.
Claim(s) 8 and 10 is rejected under 35 U.S.C. 101 because the claimed invention is directed to non-statutory subject matter. The claim(s) do not fall within at least one of the four categories of patent eligible subject matter because they cover both statutory and non-statutory embodiments (under the broadest reasonable interpretation of the claim when read in light of the specification and in view of one skilled in the art) and embraces subject matter that is not eligible for patent protection and therefore is directed to non-statutory subject matter.
“[a] transitory, propagating signal … is not a “process, machine, manufacture, or composition of matter.” Those four categories define the explicit scope and reach of subject matter patentable under 35 U.S.C. § 101; thus, such a signal cannot be patentable subject matter.” (In re Petrus A.C.M. Nuijten; Fed Cir, 2006-1371, 9/20/2007).
Specifically, Applicant’s specification describes at paragraph ¶ [0086] of the specification recites: “Some embodiments of the present application further provide a computer-readable storage medium on which a computer program is stored” Paragraph ¶ [0087] of the specification recites: “Some embodiments of the present application further provide a computer program product, the computer program product comprising a computer program” describes and as a result is drawn to a recording medium that covers both transitory and non-transitory embodiments. Thus, the claims are not eligible subject matter. It is recommended to amend and narrow the claims to cover only statutory embodiments to avoid a rejection under 35 U.S.C. § 101 by adding the limitation "non-transitory" to the claims.
Claim Rejections - 35 USC § 112
The following is a quotation of 35 U.S.C. 112(b):
(b) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention.
The following is a quotation of 35 U.S.C. 112 (pre-AIA ), second paragraph:
The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the applicant regards as his invention.
Claim 1, 3 and 7 rejected under 35 U.S.C. 112(b) or 35 U.S.C. 112 (pre-AIA ), second paragraph, as being indefinite for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor (or for applications subject to pre-AIA 35 U.S.C. 112, the applicant), regards as the invention.
Claim(s) 1 recites the limitation "prediction result" as input into the model and as output of the model in Pg. 1, ln. 11. According Fig. 5 and 2 the input is sample images and annotations. But because the claim is calling it prediction result when the output of the model is also called prediction result. It’s indefinite what the prediction results are. The interpretation for the term “prediction result” for the examiner is the output of the model. For the purpose of examination the examiner is interpreting the input prediction result as the image/sample.
Claim(s) 3 recites the limitation "prediction result" as input into the model in in Pg. 1, ln. 29 and 31. According Fig. 5 and 2 the input is sample images and annotations. But because the claim is calling it prediction result when the output of the model is also called prediction results. It’s indefinite what the prediction results are. The interpretation for the term “prediction result” for the examiner is the output of the model. For the purpose of examination the examiner is interpreting the input prediction result as the image/sample.
Claim(s) 7 recites the limitation "prediction result" as input into the model and as output of the model in in Pg. 2, ln. 33 and Pg. 3, ln. 1 and 3. According Fig. 5 and 2 the input is sample images and annotations. But because the claim is calling it prediction result when the output of the model is also called prediction result. It’s indefinite what the prediction results are. The interpretation for the term “prediction result” for the examiner is the output of the model. For the purpose of examination the examiner is interpreting the input prediction result as the image/sample.
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows:
1. Determining the scope and contents of the prior art.
2. Ascertaining the differences between the prior art and the claims at issue.
3. Resolving the level of ordinary skill in the pertinent art.
4. Considering objective evidence present in the application indicating obviousness or non-obviousness.
Claim(s) 1, 4, 6–10 are rejected under 35 U.S.C. 103 as being unpatentable over Liu et al. (See NPL attached, "Learning a Rotation Invariant Detector with Rotatable Bounding Box", hereafter, "Liu") in view of Wang et al. (US 20220020175 A1, hereafter, "Wang").
Regarding claim 1, Liu teaches A method for training a rotated bounding box object detection model (See Liu, [Abstract], In this article, a new detection method is proposed which applies the newly defined rotatable bounding box (RBox)), characterized in that it comprises:
[constructing a training dataset, wherein the training dataset comprises:
prediction result and annotation result of rotated bounding boxes in the prediction result, the annotation result comprising pixel coordinates and rotation angles of the rotated bounding boxes;
inputting the prediction result into the rotated bounding box object detection model to be trained to obtain prediction result, wherein the prediction result comprises: a predicted pixel coordinate value and a predicted angle value]; and
by comparing the prediction result with the annotation result, using the loss function to optimize the rotated bounding box object detection model to be trained, to obtain a trained rotated bounding box object detection model (See Liu, [Pg. 4, 3.2. Training, Col. 2, ln. 10–13], The RBox regression loss
L
r
b
o
x
(
x
,
l
,
g
)
is similar to SSD and Faster R-CNN, where we calculate the smooth L1 loss between the predicted RBox l and the ground truth RBox g. [Pg. 4, Col. 2, ln. 20–25], Equations 6a , 6b and 6c are the location regression terms, the size regression terms and the angle regression term, respectively).
However, Liu fail(s) to teach constructing a training dataset, wherein the training dataset comprises: prediction result and annotation result of rotated bounding boxes in the prediction result, the annotation result comprising pixel coordinates and rotation angles of the rotated bounding boxes; inputting the prediction result into the rotated bounding box object detection model to be trained to obtain prediction result, wherein the prediction result comprises: a predicted pixel coordinate value and a predicted angle value.
Wang, working in the same field of endeavor, teaches: constructing a training dataset, wherein the training dataset (See Wang, ¶ [0022], Step S101, obtaining training sample data including a first remote sensing image and position annotation information of an anchor box of a subject to be detected in the first remote sensing image, where the position annotation information including angle information of the anchor box relative to a preset direction) comprises:
prediction result and annotation result of rotated bounding boxes in the prediction result, the annotation result comprising pixel coordinates and rotation angles of the rotated bounding boxes (See Wang, ¶ [0022], Step S101, obtaining training sample data including a first remote sensing image and position annotation information of an anchor box of a subject to be detected in the first remote sensing image, where the position annotation information including angle information of the anchor box relative to a preset direction. Note: the position annotation is pixel coordinates since it's relative to the image coordinates);
inputting the prediction result into the rotated bounding box object detection model to be trained to obtain prediction result, wherein the prediction result comprises: a predicted pixel coordinate value and a predicted angle value (See Wang, ¶ [0032], Step S102, obtaining an object feature map of the first remote sensing image based on an object detection model, performing object detection on the subject to be detected based on the object feature map to obtain an object bounding box. ¶ [0039], The parameter information may include one of or any combination of a length, a width, coordinates of a center point and an angle of the bounding box).
Thus, it would have been obvious to one of ordinary skills in the art before the effective filing date of the claimed invention to modify Liu’s reference to constructing a training dataset, wherein the training dataset comprises: prediction result and annotation result of rotated bounding boxes in the prediction result, the annotation result comprising pixel coordinates and rotation angles of the rotated bounding boxes; inputting the prediction result into the rotated bounding box object detection model to be trained to obtain prediction result, wherein the prediction result comprises: a predicted pixel coordinate value and a predicted angle value based on the method of Wang’s reference. The suggestion/motivation would have been to improve training and provide more accurate position calibration of the subject (See Wang, ¶ [0028]).
Further, one skilled in the art could have combined the elements as described above by known method with no change in their respective functions, and the combination would have yielded nothing more than predictable results.
Therefore, it would have been obvious to combine Wang with Liu to obtain the invention as specified in claim 1.
Regarding claim 4, Liu teaches The method according to claim 1, characterized in that the comparing of the prediction result with the annotation result, using the loss function to optimize the rotated bounding box object detection model to be trained, to obtain the trained rotated bounding box object detection model comprises (See Liu, [Pg. 4, 3.2. Training, Col. 2, ln. 10–13], The RBox regression loss
L
r
b
o
x
(
x
,
l
,
g
)
is similar to SSD and Faster R-CNN, where we calculate the smooth L1 loss between the predicted RBox l and the ground truth RBox g. [Pg. 4, Col. 2, ln. 20-25], Equations 6a , 6b and 6c are the location regression terms, the size regression terms and the angle regression term, respectively):
using the loss function to calculate a loss between the prediction result and annotation result to obtain a pixel coordinate loss value and angle loss value (See Liu, [Pg. 4, 3.2. Training, Col. 2, ln. 10–13], The RBox regression loss
L
r
b
o
x
(
x
,
l
,
g
)
is similar to SSD and Faster R-CNN, where we calculate the smooth L1 loss between the predicted RBox l and the ground truth RBox g. [Pg. 4, Col. 2, ln. 20-25], Equations 6a , 6b and 6c are the location regression terms, the size regression terms and the angle regression term, respectively); and
using the pixel coordinate loss value and the angle loss value to adjust parameters of the rotated bounding box object detection model to be trained, to obtain the trained rotated bounding box object detection model (See Liu, [Pg. 4, 3.2. Training, Col. 2, ln. 10–13], The RBox regression loss
L
r
b
o
x
(
x
,
l
,
g
)
is similar to SSD and Faster R-CNN, where we calculate the smooth L1 loss between the predicted RBox l and the ground truth RBox g. [Pg. 4, Col. 2, ln. 20-25], Equations 6a , 6b and 6c are the location regression terms, the size regression terms and the angle regression term, respectively. [Pg. 4, Col. 2, ln. 27–28], The minimization of the angle regression term ensures that the correct angle is learned during training).
Regarding claim 6, Liu in view of Wang teaches a method for rotated bounding box object detection (See Liu, [Abstract], In this article, a new detection method is proposed which applies the newly defined rotatable bounding box (RBox). The proposed detector (DRBox) can effectively handle the situation where the orientation angles of the objects are arbitrary), characterized in that it comprises:
[obtaining a rotated bounding box annotation image to be detected]; and
inputting the rotated bounding box annotation image to be detected into the trained rotated bounding box object detection model obtained from the method as claimed in claim 1, and obtaining a rotated bounding box object detection result (See Liu, [Pg. 3, Col. 2, ln. 33–36], DRBox uses a convolutional structure for detection, as shown in Figure 1. The input image goes through multi-layer convolution networks to generate detection results).
However, Liu fail(s) to teach obtaining a rotated bounding box annotation image to be detected.
Wang, working in the same field of endeavor, teaches: obtaining a rotated bounding box annotation image to be detected (See Wang, ¶ [0022], Step S101, obtaining training sample data including a first remote sensing image and position annotation information of an anchor box of a subject to be detected in the first remote sensing image, where the position annotation information including angle information of the anchor box relative to a preset direction. Note: the position annotation is pixel coordinates since it's relative to the image coordinates).
Thus, it would have been obvious to one of ordinary skills in the art before the effective filing date of the claimed invention to modify Liu’s reference to obtaining a rotated bounding box annotation image to be detected based on the method of Wang’s reference. The suggestion/motivation would have been to improve training and provide more accurate position calibration of the subject (See Wang, ¶ [0028]).
Further, one skilled in the art could have combined the elements as described above by known method with no change in their respective functions, and the combination would have yielded nothing more than predictable results.
Therefore, it would have been obvious to combine Wang with Liu to obtain the invention as specified in claim 6.
Regarding claim 7, Liu teaches a device for training a rotated bounding box object detection model (See Liu, [Abstract], In this article, a new detection method is proposed which applies the newly defined rotatable bounding box (RBox). The proposed detector (DRBox) can effectively handle the situation where the orientation angles of the objects are arbitrary), characterized in that it comprises:
[a construction module, which is used for constructing a training dataset, wherein the training dataset comprises:
a prediction result and an annotation result of a rotated bounding box in the prediction result, the annotation result comprising pixel coordinates and a rotation angle of the rotated bounding box;
a prediction module, which is used for inputting the prediction result into a rotated bounding box object detection model to be trained, and obtaining a prediction result, wherein the prediction result comprises:
a predicted pixel coordinate value and a predicted angle value]; and
a training module, which is used for comparing the prediction result with the annotation result, using a loss function to optimize the rotated bounding box object detection model to be trained, and obtaining a trained rotated bounding box object detection model (See Liu, [Pg. 4, 3.2. Training, Col. 2, ln. 10–13], The RBox regression loss
L
r
b
o
x
(
x
,
l
,
g
)
is similar to SSD and Faster R-CNN, where we calculate the smooth L1 loss between the predicted RBox l and the ground truth RBox g. [Pg. 4, Col. 2, ln. 20–25], Equations 6a , 6b and 6c are the location regression terms, the size regression terms and the angle regression term, respectively. [Pg. 4, Col. 2, ln. 27–28], The minimization of the angle regression term ensures that the correct angle is learned during training).
However, Liu fail(s) to teach a construction module, which is used for constructing a training dataset, wherein the training dataset comprises: a prediction result and an annotation result of a rotated bounding box in the prediction result, the annotation result comprising pixel coordinates and a rotation angle of the rotated bounding box; a prediction module, which is used for inputting the prediction result into a rotated bounding box object detection model to be trained, and obtaining a prediction result, wherein the prediction result comprises: a predicted pixel coordinate value and a predicted angle value.
Wang, working in the same field of endeavor, teaches: a construction module, which is used for constructing a training dataset (See Wang, ¶ [0022], Step S101, obtaining training sample data including a first remote sensing image and position annotation information of an anchor box of a subject to be detected in the first remote sensing image, where the position annotation information including angle information of the anchor box relative to a preset direction), wherein the training dataset comprises:
a prediction result and an annotation result of a rotated bounding box in the prediction result, the annotation result comprising pixel coordinates and a rotation angle of the rotated bounding box (See Wang, ¶ [0022], Step S101, obtaining training sample data including a first remote sensing image and position annotation information of an anchor box of a subject to be detected in the first remote sensing image, where the position annotation information including angle information of the anchor box relative to a preset direction. Note: the position annotation is pixel coordinates since it's relative to the image coordinates);
a prediction module, which is used for inputting the prediction result into a rotated bounding box object detection model to be trained, and obtaining a prediction result, wherein the prediction result (See Wang, ¶ [0032], Step S102, obtaining an object feature map of the first remote sensing image based on an object detection model, performing object detection on the subject to be detected based on the object feature map to obtain an object bounding box. ¶ [0039], The parameter information may include one of or any combination of a length, a width, coordinates of a center point and an angle of the bounding box) comprises:
a predicted pixel coordinate value and a predicted angle value (See Wang, ¶ [0039], In a detection process of the object bounding box, such a technique as region of interest may be used in the classification and regression sub-network to predict and obtain, based on the object feature map, multiple bounding boxes of the subject to be detected, as well as parameter information of the obtained multiple bounding boxes. The parameter information may include one of or any combination of a length, a width, coordinates of a center point and an angle of the bounding box).
Thus, it would have been obvious to one of ordinary skills in the art before the effective filing date of the claimed invention to modify Liu’s reference to a construction module, which is used for constructing a training dataset, wherein the training dataset comprises: a prediction result and an annotation result of a rotated bounding box in the prediction result, the annotation result comprising pixel coordinates and a rotation angle of the rotated bounding box; a prediction module, which is used for inputting the prediction result into a rotated bounding box object detection model to be trained, and obtaining a prediction result, wherein the prediction result comprises: a predicted pixel coordinate value and a predicted angle value based on the method of Wang’s reference. The suggestion/motivation would have been to improve training and provide more accurate position calibration of the subject (See Wang, ¶ [0028]).
Further, one skilled in the art could have combined the elements as described above by known method with no change in their respective functions, and the combination would have yielded nothing more than predictable results.
Therefore, it would have been obvious to combine Wang with Liu to obtain the invention as specified in claim 7.
Regarding claim 8, Liu in view of Wang teaches [a computer-readable storage medium, characterized in that the computer-readable storage medium stores a computer program, wherein the computer program is executed by a processor to perform] the method as claimed in claim 1.
However, Liu fail(s) to teach a computer-readable storage medium, characterized in that the computer-readable storage medium stores a computer program, wherein the computer program is executed by a processor to perform.
Wang, working in the same field of endeavor, teaches: a computer-readable storage medium, characterized in that the computer-readable storage medium stores a computer program, wherein the computer program is executed by a processor to perform (See Wang, ¶ [0010], According to a fifth aspect of the present disclosure, an electronic device is provided, including: at least one processor; and a memory in communication connection with the at least one processor).
Thus, it would have been obvious to one of ordinary skills in the art before the effective filing date of the claimed invention to modify Liu’s reference to a computer-readable storage medium, characterized in that the computer-readable storage medium stores a computer program, wherein the computer program is executed by a processor to perform based on the method of Wang’s reference. The suggestion/motivation would have been to improve training and provide more accurate position calibration of the subject (See Wang, ¶ [0028]).
Further, one skilled in the art could have combined the elements as described above by known method with no change in their respective functions, and the combination would have yielded nothing more than predictable results.
Therefore, it would have been obvious to combine Wang with Liu to obtain the invention as specified in claim 8.
Regarding claim 9, Liu in view of Wang teaches [an electronic device, characterized in that it comprises a memory, a processor, and a computer program stored on the memory and executed on the processor, wherein the computer program when executed by the processor] performs the method as claimed in claim 1.
However, Liu fail(s) to teach an electronic device, characterized in that it comprises a memory, a processor, and a computer program stored on the memory and executed on the processor, wherein the computer program when executed by the processor.
Wang, working in the same field of endeavor, teaches: an electronic device, characterized in that it comprises a memory, a processor, and a computer program stored on the memory and executed on the processor, wherein the computer program when executed by the processor (See Wang, ¶ [0010], According to a fifth aspect of the present disclosure, an electronic device is provided, including: at least one processor; and a memory in communication connection with the at least one processor).
Thus, it would have been obvious to one of ordinary skills in the art before the effective filing date of the claimed invention to modify Liu’s reference to an electronic device, characterized in that it comprises a memory, a processor, and a computer program stored on the memory and executed on the processor, wherein the computer program when executed by the processor based on the method of Wang’s reference. The suggestion/motivation would have been to improve training and provide more accurate position calibration of the subject (See Wang, ¶ [0028]).
Further, one skilled in the art could have combined the elements as described above by known method with no change in their respective functions, and the combination would have yielded nothing more than predictable results.
Therefore, it would have been obvious to combine Wang with Liu to obtain the invention as specified in claim 9.
Regarding claim 10, Liu in view of Wang teaches [a computer program product, characterized in that the computer program product comprises a computer program, wherein the computer program, when executed by a processor], executes the method as claimed in claim 1.
However, Liu fail(s) to teach a computer program product, characterized in that the computer program product comprises a computer program, wherein the computer program, when executed by a processor.
Wang, working in the same field of endeavor, teaches: a computer program product, characterized in that the computer program product comprises a computer program, wherein the computer program, when executed by a processor (See Wang, ¶ [0010], According to a fifth aspect of the present disclosure, an electronic device is provided, including: at least one processor; and a memory in communication connection with the at least one processor).
Thus, it would have been obvious to one of ordinary skills in the art before the effective filing date of the claimed invention to modify Liu’s reference to a computer program product, characterized in that the computer program product comprises a computer program, wherein the computer program, when executed by a processor based on the method of Wang’s reference. The suggestion/motivation would have been to improve training and provide more accurate position calibration of the subject (See Wang, ¶ [0028]).
Further, one skilled in the art could have combined the elements as described above by known method with no change in their respective functions, and the combination would have yielded nothing more than predictable results.
Therefore, it would have been obvious to combine Wang with Liu to obtain the invention as specified in claim 10.
Claim(s) 2 is rejected under 35 U.S.C. 103 as being unpatentable over Liu et al. (See NPL attached, "Learning a Rotation Invariant Detector with Rotatable Bounding Box", hereafter, "Liu") in view of Wang et al. (US 20220020175 A1, hereafter, "Wang") further in view of Huahui et al. (CN118587680 A, hereafter, "Huahui").
Regarding claim 2, Liu in view Wang teaches the method according to claim 1, characterized in that the rotated bounding box object detection model to be trained (See Liu, [Abstract], In this article, a new detection method is proposed which applies the newly defined rotatable bounding box (RBox). The proposed detector (DRBox) can effectively handle the situation where the orientation angles of the objects are arbitrary) further comprises:
[a data preprocessing module, an image feature extraction module, a text feature extraction module, a feature enhancement module, and a model output module;
wherein the model output module comprises: a pixel coordinate output dimension and an angle output dimension].
However, Liu fail(s) to teach a data preprocessing module, an image feature extraction module, a text feature extraction module, a feature enhancement module, and a model output module; wherein the model output module comprises: a pixel coordinate output dimension and an angle output dimension.
Wang, working in the same field of endeavor, teaches: a data preprocessing module (See Wang, ¶ [0061], During data preprocessing, the coordinate sequence of the four vertices may be used in calculation to obtain the position annotation information of the anchor box, including the coordinates of the center point, the length, the width and the angle information of the anchor box, which will be inputted to the object detection model for model training. Note: Examiner is interpreting the preprocessing as the preprocessing module);
a model output module (See Wang, ¶ [0032], Step S102, obtaining an object feature map of the first remote sensing image based on an object detection model, performing object detection on the subject to be detected based on the object feature map to obtain an object bounding box. ¶ [0039], The parameter information may include one of or any combination of a length, a width, coordinates of a center point and an angle of the bounding box);
wherein the model output module comprises: a pixel coordinate output dimension and an angle output dimension (See Wang, ¶ [0032], Step S102, obtaining an object feature map of the first remote sensing image based on an object detection model, performing object detection on the subject to be detected based on the object feature map to obtain an object bounding box. ¶ [0039], The parameter information may include one of or any combination of a length, a width, coordinates of a center point and an angle of the bounding box).
Thus, it would have been obvious to one of ordinary skills in the art before the effective filing date of the claimed invention to modify Liu’s reference to a data preprocessing module, an image feature extraction module, a text feature extraction module, a feature enhancement module, and a model output module; wherein the model output module comprises: a pixel coordinate output dimension and an angle output dimension based on the method of Wang’s reference. The suggestion/motivation would have been to improve training and provide more accurate position calibration of the subject (See Wang, ¶ [0028]).
However, Liu and Wang fail(s) to teach an image feature extraction module, a text feature extraction module, a feature enhancement module.
Huahui, working in the same field of endeavor, teaches: an image feature extraction module (See Huahui, ¶ [n0033], S2, obtain the 2D image features, 3D coordinate encoding features, image text description text features generated based on a large language model. Note: Examiner is interpreting the large language model as the text feature extractor and the image feature extractor), a text feature extraction module (See Huahui, ¶ [n0033], S2, obtain the 2D image features, 3D coordinate encoding features, image text description text features generated based on a large language model. Note: Examiner is interpreting the large language model as the text feature extractor and the image feature extractor), a feature enhancement module (See Huahui, ¶ [n0094], Specifically, 2D image features and text features are input, and the two features are fused through the feature enhancement module to obtain updated text features and image features).
Thus, it would have been obvious to one of ordinary skills in the art before the effective filing date of the claimed invention to modify Liu’s reference to an image feature extraction module, a text feature extraction module, a feature enhancement module based on the method of Huahui’s reference. The suggestion/motivation would have been to improve the perception of objects (See Huahui, ¶ [n0134]).
Further, one skilled in the art could have combined the elements as described above by known method with no change in their respective functions, and the combination would have yielded nothing more than predictable results.
Therefore, it would have been obvious to combine Wang and Huahui with Liu to obtain the invention as specified in claim 2.
Claim(s) 3 is rejected under 35 U.S.C. 103 as being unpatentable over Liu et al. (See NPL attached, "Learning a Rotation Invariant Detector with Rotatable Bounding Box", hereafter, "Liu") in view of Wang et al. (US 20220020175 A1, hereafter, "Wang") further in view of Huahui et al. (CN 118587680 A, hereafter, "Huahui") and further in view of Orlowski et al. (US 20190004166 A1, hereafter, "Orlowski").
Regarding claim 3, Liu in view of Wang further in view of Huahui teaches the method according to claim 1, characterized in that the inputting of the prediction result into the rotated bounding box object detection model to be trained to obtain the prediction result (See Liu, [Abstract], In this article, a new detection method is proposed which applies the newly defined rotatable bounding box (RBox). The proposed detector (DRBox) can effectively handle the situation where the orientation angles of the objects are arbitrary) comprises:
[when the prediction result is rotated, the preprocessing module updates the rotation angle;
the image feature extraction module extracts image features from the prediction result to obtain a first feature;
the text feature extraction module extracts text features from the prediction result to obtain a second feature;
the feature enhancement module enhances the first feature and the second feature to obtain an enhanced feature; and
the model output module performs category prediction on the enhanced feature to obtain the prediction result].
However, Liu and Wang fail(s) to teach the image feature extraction module extracts image features from the prediction result to obtain a first feature; the text feature extraction module extracts text features from the prediction result to obtain a second feature; the feature enhancement module enhances the first feature and the second feature to obtain an enhanced feature; and the model output module performs category prediction on the enhanced feature to obtain the prediction result.
Huahui, working in the same field of endeavor, teaches: the image feature extraction module extracts image features from the prediction result to obtain a first feature (See Huahui, ¶ [n0034], Obtain the 2D image features, 3D coordinate encoding features, and text features and text position encoding features of the image text description);
the text feature extraction module extracts text features from the prediction result to obtain a second feature (See Huahui, ¶ [n0034], Obtain the 2D image features, 3D coordinate encoding features, and text features and text position encoding features of the image text description);
the feature enhancement module enhances the first feature and the second feature to obtain an enhanced feature (See Huahui, ¶ [n0094], Specifically, 2D image features and text features are input, and the two features are fused through the feature enhancement module to obtain updated text features and image features); and
the model output module performs category prediction on the enhanced feature to obtain the prediction result (See Huahui, ¶ [n0108], When performing 3D object detection, the input consists of fused image features and image query features, which are then processed by the image decoder Image Decoder to obtain the processed 3D detection result).
Thus, it would have been obvious to one of ordinary skills in the art before the effective filing date of the claimed invention to modify Liu’s reference to the image feature extraction module extracts image features from the prediction result to obtain a first feature; the text feature extraction module extracts text features from the prediction result to obtain a second feature; the feature enhancement module enhances the first feature and the second feature to obtain an enhanced feature; and the model output module performs category prediction on the enhanced feature to obtain the prediction result based on the method of Huahui’s reference. The suggestion/motivation would have been to improve the perception of objects (See Huahui, ¶ [n0134]).
However, Liu, Wang and Huahui fail(s) to teach when the prediction result is rotated, the preprocessing module updates the rotation angle.
Orlowski, working in the same field of endeavor, teaches: when the prediction result is rotated, the preprocessing module updates the rotation angle (See Orlowski, ¶ [0098], This angle Δ is used to reformulate a revised bounding box 18 as shown in FIG. 9 for Δ1. Here a new rectangular boundary box 18 is formulated from the results of the previous step where the edges of the previous boundary box are rotated about the adjustment angle).
Thus, it would have been obvious to one of ordinary skills in the art before the effective filing date of the claimed invention to modify Liu’s reference to when the prediction result is rotated, the preprocessing module updates the rotation angle based on the method of Orlowski’s reference. The suggestion/motivation would have been to adaptively update and refine orientations of the objection (See Orlowski, ¶ [0005–0006]).
Further, one skilled in the art could have combined the elements as described above by known method with no change in their respective functions, and the combination would have yielded nothing more than predictable results.
Therefore, it would have been obvious to combine Huahui and Orlowski with Liu to obtain the invention as specified in claim 3.
Claim(s) 5 is rejected under 35 U.S.C. 103 as being unpatentable over Liu et al. (See NPL attached, "Learning a Rotation Invariant Detector with Rotatable Bounding Box", hereafter, "Liu") in view of Wang et al. (US 20220020175 A1, hereafter, "Wang") further in view of Finch et al. (US 20200364920 A1, hereafter, "Finch").
Regarding claim 5, Liu and Wang teaches the method according to claim 1, characterized in that before obtaining the trained rotated bounding box object detection model (See Liu, [Abstract], In this article, a new detection method is proposed which applies the newly defined rotatable bounding box (RBox). The proposed detector (DRBox) can effectively handle the situation where the orientation angles of the objects are arbitrary), the method further comprises:
[using a validation dataset to validate the trained rotated bounding box object detection model, to obtain a model accuracy value; and
confirming that the model accuracy value is greater than or equal to a preset threshold].
However, Liu and Wang fail(s) to teach using a validation dataset to validate the trained rotated bounding box object detection model, to obtain a model accuracy value; and confirming that the model accuracy value is greater than or equal to a preset threshold.
Finch, working in the same field of endeavor, teaches: using a validation dataset to validate the trained rotated bounding box object detection model, to obtain a model accuracy value (See Finch, ¶ [0029], The training device 110 may divide the set of images into multiple groups to train the object detection system. For example, the training device 110 may divide the set of images into a training data set and a validation data set. In this case, the training device 110 may train the object detection system using the training data set and may validate that the object detection system detects objects with a threshold accuracy using the validation data set); and
confirming that the model accuracy value is greater than or equal to a preset threshold (See Finch, ¶ [0029], The training device 110 may divide the set of images into multiple groups to train the object detection system. For example, the training device 110 may divide the set of images into a training data set and a validation data set. In this case, the training device 110 may train the object detection system using the training data set and may validate that the object detection system detects objects with a threshold accuracy using the validation data set).
Thus, it would have been obvious to one of ordinary skills in the art before the effective filing date of the claimed invention to modify Liu’s reference to using a validation dataset to validate the trained rotated bounding box object detection model, to obtain a model accuracy value; and confirming that the model accuracy value is greater than or equal to a preset threshold based on the method of Finch’s reference. The suggestion/motivation would have been to verify that objection detection model is detecting at a particular accuracy and provide accurate detections (See Finch, ¶ [0029]).
Further, one skilled in the art could have combined the elements as described above by known method with no change in their respective functions, and the combination would have yielded nothing more than predictable results.
Therefore, it would have been obvious to combine Finch with Liu and Wang to obtain the invention as specified in claim 5.
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure.
Song et al. (US 20220301258 A1) teaches disclosed in the present invention is An improved rotated rectangular bounding box annotation method, used for taking as anchor boxes samples annotation and bounding box output at predicting of a target detection and tracking algorithm
Yin et al. (US 20250239041 A1) teaches at least one processor determines bounding boxes in an image using a neural network. At least one processor extracts an initial bounding box for an object in the image. At least one processor selects a cluster center of the bounding boxes and circumscribes a cluster box around the cluster center. At least one processor rotates the initial bounding box and contents of the initial bounding box according to a rotation angle determined based on the cluster box. At least one processor detects the object in the rotated initial bounding box using a trained machine learning model.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to DION J SATCHER whose telephone number is (703)756-5849. The examiner can normally be reached Monday - Thursday 5:30 am - 2:30 pm, Friday 5:30 am - 9:30 am PST.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Henok Shiferaw can be reached at (571) 272-4637. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/DION J SATCHER/Patent Examiner, Art Unit 2676
/Henok Shiferaw/Supervisory Patent Examiner, Art Unit 2676