DETAILED ACTION
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Claims 1-12 are pending.
Claim Rejections - 35 USC § 103
The following is a quotation of pre-AIA 35 U.S.C. 103(a) which forms the basis for all obviousness rejections set forth in this Office action:
(a) A patent may not be obtained though the invention is not identically disclosed or described as set forth in section 102 of this title, if the differences between the subject matter sought to be patented and the prior art are such that the subject matter as a whole would have been obvious at the time the invention was made to a person having ordinary skill in the art to which said subject matter pertains. Patentability shall not be negatived by the manner in which the invention was made.
Claim(s) 1-6 and 8-12 is/are rejected under 35 U.S.C. 103 as being unpatentable over Sivalingam et al (US20200160106A1) in view of Masuda (US20210192680A1).
Regarding claims 1, 11 and 12, Sivalingam teaches an image processing method, by a computer, comprising:
acquiring an image made by an omnidirectional imaging;
(Sivalingam, "images acquired from cameras with large field-of-view lenses can suffer from significant geometric distortions”, [0024]; “ultra-low power sensor including... a wide (e.g., a 90° diagonal) field-of-view lens", [0025]; Masuda, "The spherical content is generated by being shot with a spherical camera capable of shooting in 360° in all directions.", [0027]; Sivalingam teaches acquiring images through a wide FOV lens (90°+). Masuda teaches acquiring 360° omnidirectional spherical content. Together they teach acquiring an omnidirectional image)
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention was made to incorporate the teachings of Masuda into the system or method of Sivalingam in order to systematically generate projection-matched distorted training images (via pitch/roll/yaw-varied equirectangular projection of face images) that replicate the target omnidirectional camera's exact distortion profile, including severe pole/edge distortion, so the detector learns to maintain accuracy where Sivalingam's conventional models "perform poorly" ([0026]). The combination of Sivalingam and Masuda also teaches other enhanced capabilities.
The combination of Sivalingam and Masuda further teaches:
executing an object detection process of detecting an object in the acquired image;
(Sivalingam, "detecting one or more objects of interest directly from the captured images", [0005]; Masuda, "The image processing apparatus 100 inputs the accepted spherical content 50 into the detection model 150... detects each face 60", [0038]; Sivalingam teaches applying a detection algorithm on captured images. Masuda teaches inputting spherical content into a model to detect faces. Together they teach executing an object detection process)
calculating a detection accuracy of the object in the object detection process;
(Sivalingam, "the object detection models perform poorly when an object of interest is located towards the edges of the image, where the geometric distortions are severe.", [0026]; Masuda, "performs learning with the learning data 145... generates the detection model 150:, [0036]; “Note that the correct label may include attribute information", [0048]; Sivalingam notes poor detection accuracy at distorted edges. Masuda teaches learning with labeled data, which inherently evaluates accuracy during training. Together they teach calculating detection accuracy)
processing the image so as to increase a distortion of the object included in the image on the basis of the detection accuracy; and
(Sivalingam, "To address one or more issues... it is proposed to use distorted images during training", [0033]; Masuda, "generates images with the equirectangular projection scheme, the images having respective angles different... referred to as a “distorted face image” because the face is distorted", [0035]; "for an object present at the pitch angle of near 90°... and normally difficult to recognize as a human face, the learning-data creation unit 132 can create learning data", [0089]; Sivalingam identifies poor accuracy at high-distortion regions and proposes training on purposefully distorted images. Masuda specifically creates distorted images by varying projection angles to learn objects "difficult to recognize". Thus, they teach processing images to increase object distortion based on detection accuracy)
outputting a processed image resulting from the processing.
(Sivalingam, "the training processor 250 may obtain the distorted training images 371", [0058]; Masuda, "The projection transformation unit 132B stores the projection-transformed image data (learning data) in the learning-data storage unit 122", [0109]; Sivalingam provides generated distorted images for model training. Masuda stores projection-transformed (distorted) images. Both teach outputting the processed image)
Regarding claim 2, the combination of Sivalingam and Masuda teaches its/their respective base claim(s).
The combination further teaches the image processing method according to claim 1, wherein
the image is an omnidirectional image having an object associated with a truth label,
(Sivalingam, "train an object detection model 173 to detect specific objects of interest", [0031]; Masuda, "image-data storage unit 121 has items such as... “part information””, [0049]; “Note that the correct label may include attribute information", [0048]; Sivalingam trains models for specific objects. Masuda provides spherical content containing objects with a "correct label" (truth label) and part information)
the detection accuracy is calculated on the basis of the truth label, and
(Masuda, "on the basis of the distorted face image included in the learning data 145, learns each feature amount... and generates a detection model", [0113]; generating a detection model using learning data with correct labels, which inherently evaluates accuracy against truth labels during training)
the processing is executed in a case where the detection accuracy is lower than a threshold.
(Sivalingam, "object detection models perform poorly when an object of interest is located towards the edges... where the geometric distortions are severe.", [0026]; Masuda, "for an object present at the pitch angle of near 90°... and normally difficult to recognize", [0089]; Both create distorted training images specifically to address scenarios where baseline models perform poorly (accuracy is below a threshold))
Regarding claim 3, the combination of Sivalingam and Masuda teaches its/their respective base claim(s).
The combination further teaches the image processing method according to claim 1, wherein
the image includes a first image and a second image different from the first image,
(Sivalingam, "training process 370 may involve training on distorted training images 371", [0045]; Masuda, "transforms original image data with the projection scheme... and creates learning data", [0069]; Both teach processing an original (first) image to create a distorted/transformed (second) image)
the detection accuracy is related to a detection result obtained by inputting the first image to a learning model trained in advance to execute the object detection process, and
(Sivalingam, "Since the inputs to the detection algorithm 125 are undistorted images 124, the training images 171 are also undistorted.", [0031]; Masuda, "selects a detection model 150 corresponding to the type of the projection transformation", [0118]; Both teach that baseline accuracy relates to inputting the original/undistorted images into pre-trained models)
the processing of the image is executed to the second image.
(Sivalingam, "existing source images 421 may be distorted in a distortion algorithm 423 to obtain the synthetic distorted images 425.", [0051]; Masuda, "performs projection transformation on the one piece of image data... stores the projection-transformed image data (learning data)", [0109]; Both teach generating the second image via the distortion-increasing processing)
Regarding claim 4, the combination of Sivalingam and Masuda teaches its/their respective base claim(s).
The combination further teaches the image processing method according to claim 3, further comprising: training the learning model using the processed image.
(Sivalingam, "training an object detection model based on the one or more distorted training images", [0004]; Masuda, "model generation unit 133 generates, on the basis of the learning data... a learned model", [0096]; Both explicitly teach training detection models using the processed/distorted images)
Regarding claim 5, the combination of Sivalingam and Masuda teaches its/their respective base claim(s).
The combination further teaches the image processing method according to claim 3, wherein
the detection accuracy is calculated for each object class, and
(Sivalingam, " detect one or more specific objects of interest (e.g., faces, human body, gestures, signs, toys, cars, patterns, logos, etc.)", [0045]; Masuda, "divide the learning data 145 on the basis of... attribute information such as race, gender, and age, and may generate a plurality of detection models", [0145]; Both detect and calculate accuracy for varying object classes)
the second image includes an object of which detection accuracy is determined to be equal to or lower than a threshold in the first image.
(Sivalingam, "object detection models perform poorly when an object of interest is located towards the edges", [0026]; Masuda, "an object present at the pitch angle of near 90°... and normally difficult to recognize", [0089]; Distorted (second) images are generated for objects that suffer from poor detection accuracy in original configurations)
Regarding claim 6, the combination of Sivalingam and Masuda teaches its/their respective base claim(s).
The combination further teaches the image processing method according to claim 1, wherein the processing of the image includes changing a default viewpoint of the image to a viewpoint that is randomly set.
(Sivalingam, "skewing the display in any combination of vertical and horizontal angles", [0050]; Masuda, "changes respective angles in the pitch direction, the roll direction, and the yaw direction by 1°", [0071]; Both systematically/randomly change viewpoint angles to generate variations)
Regarding claim 8, the combination of Sivalingam and Masuda teaches its/their respective base claim(s).
The combination further teaches the image processing method according to claim 1, wherein the processing of the image includes:
determining whether the image includes an object having at least one of an aspect ratio and a size that exceeds a reference value; and
(Masuda, "determine whether the subject is in contact with the image frame (at the left and right edge portions of the image)", [0073]; checking if a subject contacts frame boundaries, serving as a size/position reference)
making more processed images in a case where a specific object exceeding the reference value is determined to be included than in a case where the specific object is determined not to be included.
(Masuda, "creates distorted face images each having an angle of 180° in the yaw direction, for the distorted face images different in angle in the pitch direction and the roll direction.", [0092]; Creating additional training images for specific objects that exceed limits (e.g., highly distorted or split at frame edges))
Regarding claim 9, the combination of Sivalingam and Masuda teaches its/their respective base claim(s).
The combination further teaches the image processing method according to claim 1, wherein
the object detection process includes a rule-based object detection process, and
(Sivalingam, "One of the computer vision algorithms... using cascades of boosted LBP (local binary pattern) classifiers", [0025]; using hand-crafted/rule-based features like LBP classifiers)
the processing of the image is executed to the image having been subjected to the object detection process.
(Sivalingam, "inference process 320... A detection algorithm 325 may be applied”, [0043]; “The training process 370 may involve training on distorted training images", [0045]; Masuda, "selects a replacement model 160 corresponding to the predetermined target detected from the input data", [0119]; Processing is performed after subjection to initial detection processes)
Regarding claim 10, the combination of Sivalingam and Masuda teaches its/their respective base claim(s).
The combination further teaches the image processing method according to claim 1, wherein
the processing of the image includes executing a viewpoint change process to increase the distortion of the object, and
(Sivalingam, "existing images may be distorted on purpose... in a distortion algorithm 423", [0051]; Masuda, "changes an angle of the subject... performs projection transformation", [0070]; Both teach changing viewpoint angles or applying algorithms to purposefully increase distortion)
the viewpoint change process includes
projecting the image onto a unit sphere;
(Masuda, "the content 70 has a spherical shape”, [0076]; “projection-transforming the content 70 with the equirectangular projection scheme", [0077]; mapping image content to/from spherical coordinates using equirectangular projection)
setting a new viewpoint depending on the projected image; and
(Masuda, "creates image data with an angle changed stepwise", [0108]; adjusting angles stepwise to generate new spherical viewpoints)
developing the projected image to a plane having the new viewpoint at a center thereof.
(Masuda, "transformed image has an equidistance on the centerline... however, the object in the image is stretched (distorted) near the spherical poles", [0077]; Equirectangular projection flattens the spherical image to a 2D plane centered at the viewpoint, increasing distortion towards the poles)
Claim(s) 7 is/are rejected under 35 U.S.C. 103 as being unpatentable over Sivalingam et al (US20200160106A1) in view of Masuda (US20210192680A1) and further in view of Onuma (US20220383586A1).
Regarding claim 7, the combination of Sivalingam and Masuda teaches its/their respective base claim(s).
The combination further teaches the image processing method according to claim 1, wherein the processing of the image includes:
specifying an interval between two bonding boxes making a longest distance therebetween among those associated with a plurality of truth labels in the image; and
(Sivalingam, "detect one or more objects of interest in distorted images", [0004]; Masuda, "label serving as a correct answer of the image", [0048]; Onuma, "a plurality of straight lines connecting the sight designation position and the plurality of object positions", [0052]; Sivalingam teaches plurality of objects; Masuda teaches truth-label per object in wide-angle image; Sivalingam/Masuda lack longest-distance interval selection among truth-label boxes; Onuma teaches intervals between objects and selecting two pieces of sight information for longest pair)
setting a viewpoint of the image to a midpoint of the interval.
(Onuma, "a single virtual viewpoint at a point (sight designation position) w % here a plurality of pieces of sight information intersect", [0052]; "one of the two objects is located at the center", [0052]; Sivalingam/Masuda lack midpoint viewpoint setting; Onuma teaches setting viewpoint at intersection/midpoint of two-object interval)
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention was made to incorporate the teachings of Onuma into the system or method of Sivalingam and Masuda in order to specify interval between two farthest truth-label boxes to enclose plurality of objects, and to set viewpoint to midpoint of longest interval to balance distortion for both objects. The combination of Sivalingam, Masuda and Onuma also teaches other enhanced capabilities.
Conclusion
Any inquiry concerning this communication or earlier communications from the examiner should be directed to JIANXUN YANG whose telephone number is (571)272-9874. The examiner can normally be reached on MON-FRI: 8AM-5PM Pacific Time.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Amandeep Saini can be reached on (571)272-3382. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272- 1000.
/JIANXUN YANG/
Primary Examiner, Art Unit 2662 7/13/2026