Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Double Patenting
The nonstatutory double patenting rejection is based on a judicially created doctrine grounded in public policy (a policy reflected in the statute) so as to prevent the unjustified or improper timewise extension of the “right to exclude” granted by a patent and to prevent possible harassment by multiple assignees. A nonstatutory double patenting rejection is appropriate where the conflicting claims are not identical, but at least one examined application claim is not patentably distinct from the reference claim(s) because the examined application claim is either anticipated by, or would have been obvious over, the reference claim(s). See, e.g., In re Berg, 140 F.3d 1428, 46 USPQ2d 1226 (Fed. Cir. 1998); In re Goodman, 11 F.3d 1046, 29 USPQ2d 2010 (Fed. Cir. 1993); In re Longi, 759 F.2d 887, 225 USPQ 645 (Fed. Cir. 1985); In re Van Ornum, 686 F.2d 937, 214 USPQ 761 (CCPA 1982); In re Vogel, 422 F.2d 438, 164 USPQ 619 (CCPA 1970); In re Thorington, 418 F.2d 528, 163 USPQ 644 (CCPA 1969).
A timely filed terminal disclaimer in compliance with 37 CFR 1.321(c) or 1.321(d) may be used to overcome an actual or provisional rejection based on nonstatutory double patenting provided the reference application or patent either is shown to be commonly owned with the examined application, or claims an invention made as a result of activities undertaken within the scope of a joint research agreement. See MPEP § 717.02 for applications subject to examination under the first inventor to file provisions of the AIA as explained in MPEP § 2159. See MPEP § 2146 et seq. for applications not subject to examination under the first inventor to file provisions of the AIA . A terminal disclaimer must be signed in compliance with 37 CFR 1.321(b).
The filing of a terminal disclaimer by itself is not a complete reply to a nonstatutory double patenting (NSDP) rejection. A complete reply requires that the terminal disclaimer be accompanied by a reply requesting reconsideration of the prior Office action. Even where the NSDP rejection is provisional the reply must be complete. See MPEP § 804, subsection I.B.1. For a reply to a non-final Office action, see 37 CFR 1.111(a). For a reply to final Office action, see 37 CFR 1.113(c). A request for reconsideration while not provided for in 37 CFR 1.113(c) may be filed after final for consideration. See MPEP §§ 706.07(e) and 714.13.
The USPTO Internet website contains terminal disclaimer forms which may be used. Please visit www.uspto.gov/patent/patents-forms. The actual filing date of the application in which the form is filed determines what form (e.g., PTO/SB/25, PTO/SB/26, PTO/AIA /25, or PTO/AIA /26) should be used. A web-based eTerminal Disclaimer may be filled out completely online using web-screens. An eTerminal Disclaimer that meets all requirements is auto-processed and approved immediately upon submission. For more information about eTerminal Disclaimers, refer to www.uspto.gov/patents/apply/applying-online/eterminal-disclaimer.
Claim 1 is rejected under the judicially created doctrine of obviousness-type double patenting as
being unpatentable over claim 1 of U.S. Patent No. 10,339,422, U.S. Patent No. 11,734,920 and U.S. Patent No. 10,572,772. The conflicting claims are not identical because patent' 422 claim 1 requires the additional elements of "one or more processors acting as a learning unit configured to learn the dictionary based on the teacher data, wherein the accepting unit accepts the selection of the partial region based on a location on the input image subjected to the one operation", not required by claim 1 of the instant application. Additionally, patent' 772 claim 1 requires the additional elements of "determining a location of the new partial region based on a start point of the second input with respect to the first image and an end point of the second input with respect to the first image; displaying, on the display device with respect to the first image, the new partial region; and generating teacher data based on the new partial region and the class " not required by claim 1 of the instant application. However, the conflicting claims are not patentably distinct from each other because:
Claims 1 of application '426 and claim 1 of patent '422, patent’ 920 and patent '772 recite common subject
matter;
Whereby claim 1, which recites the open ended transitional phrase "comprising", does not
preclude the additional elements recited by claim 1 of the patent, and
Whereby the elements of claim 1 are fully anticipated by patent claim 1, and anticipation is "the
ultimate or epitome of obviousness" (In re Kalm, 154 USPQ 10 (CCPA 1967), also In re Dailey, 178 USPQ 293 (CCPA 1973) and In re Pearson, 181 USPQ 641 (CCPA 1974)).
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claim(s) 1-20 are rejected under 35 U.S.C. 103 as being unpatentable over Kanda et al ( US 2007/009152) in view of Yoshihiko et al (JP 2014/059729).
As to claim 1, Kanda teaches the system comprising: a storage configured to store a program; and one or more processors, wherein the one or more processors are configured to execute the program to perform: detecting objects from a plurality of input images, wherein the plurality of input images includes a first image( ( Kanda [0008] teaches a learning-type classifying apparatus comprises defective region extracting unit for extracting defective regions of classification targets from an image in which the plurality of regions of the classification targets are present and teacher data creating unit for each of the regions so that the classification results of the integrated regions are reflected in each region included in the integrated regions and creating teacher data for extract defective regions of classification targets, s2, figure 2); displaying, on a display device, a graph includes a relation between the plurality of input images and the detected objects; receiving an input of a location on the graph ( Kanda [0053] teaches in such a user-interactive operation as the creation of the teacher data, a plurality of regions of classification targets in an image are displayed while being automatically integrated in a suitable manner Kanda [0043] teaches a user determines whether a judgment on the display results on the display screen is correct or wrong (step S7). When the judgment is correct, the user instructs using the input unit 109 that the judgment results are OK (step S8). After the OK instruction in step S8, the classification results of the integrated regions are considered as the classification results of the respective regions included in each of the integrated regions, thereby creating the teacher data in the teacher data creating unit 105 (step S9). FIG. 4 shows the teacher data created on the basis of the display state shown in FIG. 3), displaying, on the display device, the first image of the plurality of input images in accordance with the location on the graph ((Kanda [0039] teaches after the characteristic values of the regions are calculated, the defective regions are classified into predetermined defect kinds
in the classifying unit 103 (step S4). As a method of classification, there is a method (k neighborhood
method) comprising restoring several teacher data in the teacher data storing unit 106, selecting, from
these teacher data, k teacher data located in the neighborhood of the defective regions of the
classification targets within a characteristic space, and classifying them into a defect kind having the
largest number. It is to be noted that when there is no previous teacher data at the initial stage of
learning, defective regions in a screen are properly classified and assigned with temporary defective
kinds (information on a correct defect kind is obtained by the subsequent user operation); Kanda [0042] teaches FIG. 3 shows one example of a display screen on the display unit 108 when step S6 is executed. The screen has a display area 150 for displaying each of the integrated regions in a color-coded manner, and a display area 141 for displaying the defect kinds of the integrated regions. Here, displays of the same color are expressed with the same kind of design. Therefore, among regions 151 to 159 displayed in the display area 150 in FIG. 3, the regions 151 to 154,158 and 159 are displayed in an integrated manner in the same color as regions belonging to the same defect kind, the regions 155 and 156 are displayed in an integrated manner in another same color as regions belonging to another same defect kind, and the region 157 is displayed in still another color as a region belonging to still another same defect kind. Kanda [0053] teaches in such a user-interactive operation as the creation of the teacher data, a plurality of regions of classification targets in an image are displayed while being automatically integrated in a suitable manner, and the teacher data is automatically created so that the contents of correction for the integrated regions are reflected in the respective regions in the integrated regions, thereby making it possible to save the trouble of correcting the individual regions and to efficiently create the teacher data for the image in which a plurality of classification targets are present.
While Kanda teaches the limitation above, Kanda fails to teach " receiving a selection of a new partial region, wherein the new partial region is selected by inputting at least two points with respect to the first image, wherein the new partial region corresponds to a selected class; displaying, on the display device with respect to the first image, the new partial region; and generating data based on the new partial region and the selected class at least.”
Yoshihiko et al teaches dictionary data for object identification is prepared by mutually using different kinds of data such as comprehensively collected camera images from a plurality of places, time information at the time of image data acquisition, position data, and meteorological data, SO that dictionary data commonly applicable to object detecting devices installed in different locations is generated. By updating the dictionary data using newly collected image, recognition performance is improved continuously while the object detecting devices are operated ( abstract). Yoshihiko clearly teaches vehicle area is used as teaching data in learning. The feature extraction unit 45 extracts higher-order feature amounts from the region cut out by the object automatic detection unit 44. Specifically, a plurality of types of higher-order feature quantities are extracted from the teaching data, and higher-order feature quantities that are considered to be effective are ranked based on the label information. The learning unit 46 performs learning based on the labeled higher-order feature quantity extracted by the feature extraction unit 45, and selects a higher-order feature quantity that exhibits the highest performance. The dictionary construction unit 47 generates dictionary data based on the learning
result of the learning unit 46, note that a dictionary is constructed from videos at a plurality of
points, and can be commonly incorporated in predetermined object detection apparatuses having
different installation locations ( figure 5) the video data sorting unit 43 sorts video data necessary for learning of the discriminator generating device based on the content estimation result of the input video. Further, the result of estimating the content of the video is given to the video data as label information ( figure 7). It would have been obvious to one skilled in the art before filing of the claimed invention to modify the objects taught by Kanda with the teaching of Yoshihiko to use the object automatic detection unit 44 automatically extracts an object (vehicle) region from the arbitrary frame image 3 of the previous camera image with reference to the recognition dictionary 2 by the pattern
recognition system algorithm, and the object region The feature extraction unit 45 extracts the
higher-order feature quantity 3, the learning unit 46 learns the recognition dictionary 2, and the
dictionary construction unit 47 upgrades the dictionary data based on the learning result to
generate the recognition dictionary Thereby, since the dictionary data is updated in the best
direction, the object recognition performance can be improved continuously while the object
detecting devices are operated. Thus, the claimed subject matter would have been obvious to a person having ordinary skill in the art before the effective filing date of the claimed invention.
As to claim 2, Kanda teaches the system according to claim 1, wherein the data includes location information of the new partial region on the first image corresponding to the selected class((Kanda [0039] teaches after the characteristic values of the regions are calculated, the defective regions are classified into predetermined defect kinds in the classifying unit 103 (step S4). As a method of classification, there is a method (k neighborhood method) comprising restoring several teacher data in the teacher data storing unit 106, selecting, from these teacher data, k teacher data located in the neighborhood of the defective regions of the classification targets within a characteristic space, and classifying them into a defect kind having the largest number. It is to be noted that when there is no previous teacher data at the initial stage of learning, defective regions in a screen are properly classified and assigned with temporary defective kinds (information on a correct defect kind is obtained by the subsequent user operation).
As to claim 3, Kanda teaches the system according to claim 2, wherein the location included in the data comprises at least one coordinate value of vertex belonging to the new partial region(paragraph [0056-0057]; [0060]).
As to claim 4, Kanda teaches the system according to claim 1, wherein the data includes at least one coordinate value of vertex belonging to the new partial region on the first image, the selected class corresponding to the new partial region, and frame number of the first image (paragraph [0056-0057]; [0060]).
As to claim 5, Kanda teaches the system according to claim 4, wherein the data includes at least one coordinate value of vertex belonging to a partial region of the detected object on the first image, and a class of the detected object (paragraph [0056-0057]; [0060]).
As to claim 6, Yoshihiko et al teaches the system according to claim 1, wherein a shape of the new partial region is rectangle (figure 4) .
As to claim 7,Yoshihiko et al teaches the system according to claim 1, wherein the one or more processors are further configured to execute the program to perform: updating the first image on the display device based on receiving the new partial region (recognition dictionary 1 from an arbitrary frame image 2 of an arbitrary camera image (the object automatic detection unit refers to a vehicle is traveling on a curved road) by a pattern recognition system algorithm. The object (vehicle) region is automatically cut out, the higher-order feature quantity 2 is extracted from the object region by the feature extraction unit 45, the learning dictionary 46 learns the recognition dictionary 1, and the dictionary construction unit 47 obtains the learning result that is combined with the first frame Based on the dictionary data upgrade, the recognition dictionary 2 is generated; paragraph [0044])
As to claim 8, Kanda et al teaches the system according to claim 1, wherein the one or more processors are further configured to execute the program to perform: emphasizing, on the displayed first image, the new partial region( Kanda [0039] teaches after the characteristic values of the regions are calculated, the defective regions are classified into predetermined defect kinds in the classifying unit 103 (step S4). As a method of classification, there is a method (k neighborhood method) comprising restoring several teacher data in the teacher data storing unit 106, selecting, from these teacher data, k teacher data located in the neighborhood of the defective regions of the classification targets within a characteristic space, and classifying them into a defect kind having the largest number. It is to be noted that when there is no previous teacher data at the initial stage of learning, defective regions in a screen are properly classified and assigned with temporary defective kinds (information on a correct defect kind is obtained by the subsequent user operation).
The limitation of claims 9-20 has been addressed above.
Reference Cited
Cho, US-20220036066, discloses "Disclosed are an X-RAY image reading
support method including the steps of acquiring a target X-RAY image photographed by
transmitting or reflecting X-RAY in a reading space in which an object to be read is
disposed; applying the target X-RAY image to a reading model that extracts features
from an input image; and identifying the object to be read as an object corresponding to
a classified class when the object to be read is classified as a set class based on a first
feature set extracted from the target X-RAY image, and an X-RAY image reading
support system performing the method." (Abstract)
Schmidt, US-20080292194, discloses "A method and system for segmenting an
object represented in one or more input images, each of the one or more input images
comprising a plurality of pixels. The method comprising: aligning the one or more input
images with one or more corresponding template images each comprising a plurality of
pixels; extracting features of each of the one or more input images and one or more
template images; and classifying each pixel, or a group of pixels, in the one or more
input images based on the measured features of the one or more input images and the
one or more corresponding template images in accordance with a classification model
mapping image properties or features to a respective class so as to segment the object
represented in the one or more input images according to the classification of each pixel
or group of pixels." (Abstract)
Jones, US-8478020, discloses "A customer account number is received via an
interface of a document processing device. A plurality of documents associated with the
deposit transaction is received in an input receptacle of the document processing
device. The plurality of documents is transported, one at a time, along a transport path
from the input receptacle past an image scanner to one or more output receptacles.
Each document is imaged with the image scanner to produce image data associated
with the deposit transaction. The image data is reproducible as a visually readable
image of at least a portion of each document. Deposit information is generated from the
image data associated with the deposit transaction. The customer account number is
associated with the generated deposit information. The deposit information is
transmitted from the document processing device to a teller system." (Abstract)
Klein, US-20150146963, discloses "Currency bills are transported past an image
scanner to one or more output receptacles. Each of the bills is imaged to produce image
data from which a visually readable image of each bill can be reproduced. The serial
number, denomination, and/or secondary identifiers of a bill is attempted to be extracted
and/or determined from the image data associated with the bill. The serial number of the
bill has an integer number, X, of characters. One or more of the X characters of the
serial number of the currency bill is not extracted with a predetermined confidence. In
response to failing to extract all of the X characters of the serial number of the bill with
the predetermined confidence, a serial number field in an electronic record associated
with the bill is populated with a serial number snippet image. The electronic record is
stored in a non-transitory memory." (Abstract)
Sullivan, US-20160267382, discloses "A method of analyzing a color image is
disclosed. The color image depicts subject matter of particular interest and/or relevant
to solving a given problem. The color image is comprised of image pixels comprising
image pixel data. The method is comprised of storing hypothesis decision output
information in multi-dimensional color look-up tables; for each pixel of the color image,
using the multi-dimensional color look-up tables to produce logic decision output
information; grouping the logic decision output information statistically to produce logic
decision output information; combining logic decision output information into statistical
hypothesis decisions for the color image; and applying the statistical hypothesis
decisions to perform an action on the subject matter of the image directed to produce a
decision regarding the problem of interest. The problem of interest may be a medical,
space or oceanographic exploration, intelligence, forensic, counterfeiting, agricultural,
meteorological, seismological, or object detection problem." (Abstract)
Shen, US-20180173997, discloses "The disclosure relates to a training device
and method for an image processing device and an image processing device. The
training device is used for training first and second image processing units, comprising:
a training unit to input a first realistic image without a specific feature into the first image
processing unit to generate a first generated image with the specific feature through first
image processing, and to input a second realistic image with the specific feature into the
second image processing unit to generate a second generated image without the
specific feature through second image processing; and a classifying unit performing
classification processing to discriminate realistic and generated images, wherein the
training unit performs first training processing of training the classifying unit based on
the realistic and generated images, and performs second training processing of training
the first and second image processing units based on the training result." (Abstract)
Shechtman, US-20190251401, discloses "The present disclosure relates to an
image composite system that employs a generative adversarial network to generate
realistic composite images. For example, in one or more embodiments, the image
composite system trains a geometric prediction neural network using an adversarial
discrimination neural network to learn warp parameters that provide correct geometric
alignment of foreground objects with respect to a background image. Once trained, the
determined warp parameters provide realistic geometric corrections to foreground
objects such that the warped foreground objects appear to blend into background
images naturally when composited together." (Abstract)
Zadeh, US-20180204111, discloses "Specification covers new algorithms,
methods, and systems for: Artificial Intelligence; the first application of General-Al
(versus Specific, Vertical, or Narrow-Al) (as humans can do); addition of reasoning,
inference, and cognitive layers/engines to learning module/engine/layer; soft computing;
Information Principle; Stratification; Incremental Enlargement Principle; deeplevel/
detailed recognition, e.g., image recognition (e.g., for action, gesture, emotion,
expression, biometrics, fingerprint, tilted or partial-face. OCR, relationship, position,
pattern, and object); Big Data analytics; machine learning; crowd-sourcing;
classification; clustering; SVM; similarity measures; Enhanced Boltzmann Machines;
Enhanced Convolutional Neural Networks; optimization; search engine; ranking;
semantic web; context analysis; question-answering system; soft, fuzzy, or un-sharp
boundaries/impreciseness/ambiguities/fuzziness in class or set, e.g., for language
analysis; Natural Language Processing (NLP); Computing-with-Words (CWW); parsing;
machine translation; music, sound, speech, or speaker recognition; video search and
analysis (e.g. tracking); image annotation; image or color correction; data reliability; Number;
Z-Web; Z-Factor: rules engine; playing games; control system; autonomous
vehicles or drones; self-diagnosis and self-repair robots; system diagnosis; medical
diagnosis; genetics; drug discovery; biomedicine; data mining; event prediction;
financial forecasting (e.g., for stocks); economics; risk assessment; fraud detection
(e.g., for cryptocurrency); e-mail management; database management; indexing and
join operation; memory management; data compression; event-centric social network;
social behavior; and Image Ad and Referral Networks." (Abstract)
Xu, US-20190333219, discloses "Techniques for generating an enhanced cone beam
computed tomography (CBCT) image using a trained model are provided. A
CBCT image of a subject is received, a synthetic computed tomography (sCT) image
corresponding to the CBCT image is generated, using a generative model. The
generative model is trained in a generative adversarial network (GAN). The generative
model is further trained to process the CBCT image as an input and provide the sCT
image as an output. The sCT image is presented for medical analysis of the subject."
(Abstract)
Fu, US-20190295302, discloses "Embodiments provide methods and systems for
image generation through use of adversarial networks. An embodiment trains an image
generator comprising (i) a generator implemented with a first neural network configured
to generate a fake image based on a target segmentation, (ii) a discriminator
implemented with a second neural network configured to distinguish a real image from a
fake image and output a discrimination result as a function thereof and (iii) a segmentor
implemented with a third neural network configured to generate a segmentation from
the fake image. The training includes (i) operating the generator to output the fake
image to the discriminator and the segmentor and (ii) iteratively operating the generator,
discriminator, and segmentor during a training period, whereby the discriminator and
generator train in an adversarial relationship with each other and the generator and
segmentor train in a collaborative relationship with each other." (Abstract)
Lim, US-20200065635, discloses "A processor-implemented object detection
method is provided. The method receives an input image, generates a latent variable
that indicates a feature distribution of the input image, and detects an object in the input
image based on the generated latent variable." (Abstract)
Akhazhanov, US-20200394475, discloses "A computer-implemented method, a
computing system, and a computer program product for generating new items
compatible with given items may use data associated with a plurality of images and
random noise data associated with a random noise image to train an adversarial
network including a series of generator networks and a series of discriminator networks
corresponding to the series of generator networks by modifying, using a loss function of
the adversarial network that depends on a compatibility of the images, one or more
parameters of the series of generator networks. The series of generator networks may
generate a generated image associated with a generated item different than the given
items." (Abstract)
Shanbhag, US-20200364864, discloses "Methods and systems are provided for
generating a normative medical image from an anomalous medical image. In an
example, the method includes receiving an anomalous medical image, wherein the
anomalous medical image includes anomalous data, mapping the anomalous medical
image to a normative medical image using a trained generative network of a generative
adversarial network (GAN), wherein the anomalous data of the anomalous medical
image is mapped to normative data in the normative medical image. In some examples,
the method may further include displaying the normative medical image via a display
device, and/or utilizing the normative medical image for further image analysis tasks to
generate robust outcomes from the anomalous medical image." (Abstract)
Brown, US-20210256707, discloses "Example aspects of the present disclosure
are directed to systems and methods that enable weakly-supervised learning of
instance segmentation by applying a cut-and-paste technique to training of a generator
model included in a generative adversarial network. In particular, the present disclosure
provides a weakly-supervised approach to object instance segmentation. In some
implementations, starting with known or predicted object bounding boxes, a generator
model can learn to generate object masks by playing a game of cut-and-paste in an
adversarial learning setup." (Abstract)
Wang, US-20210019541, discloses "Systems, methods, and computer-readable
media are provided media for transferring visual attributes to images. In some
examples, a system can obtain a first image associated with a user; generate a second
image including image data from the first image modified to add a first visual attribute
transferred from one or more images or remove a second visual attribute in the image
data; compare a first set of features from the first image with a second set of features
from the second image; determine, based on a comparison result, whether the first
image and the second image match at least partially; and update a library of user
verification images to include the second image when the first image and the second
image match at least partially." (Abstract)
Park, US-20200193269, discloses "A recognizer including a shared encoder to
extract a feature of an input image in a source domain and a target domain; and a
shared decoder to classify a class of an object included in the input image based on the
feature of the input image, in the source domain and the target domain. A set of object
classes of the source domain and a set of object classes of the target domain differ from
each other." (Abstract)
Nam, US-10733733, discloses "There is provided an anomaly detection method,
apparatus, and system that can improve the accuracy and reliability of a detection result
using GAN (Generative Adversarial Networks). An anomaly detection apparatus
according to some embodiments includes a memory that stores a GAN-based image
translation model and an anomaly detection model, and a processor that translates a
learning image with a low-difficulty level into a learning image with a high-difficulty level
and learns the anomaly detection model using the translated learning image. The
anomaly detection apparatus can improve the detection performance by learning the
anomaly detection model with the learning image with the high-difficulty level in which it
is difficult detect the anomaly." (Abstract)
Zhu, US-12469137, discloses "Apparatuses, systems, and techniques to
generate one or more data items related to an object from a set of multiple data items
related to the same object using a generative adversarial network. In at least one
embodiment, data objects in a set are disentangled into common and unique
components and encoded such that they may be used to train one or more neural
networks to generate missing data from the data objects in the set, such as generating
missing medical images from a set of related medical images." (Abstract)
Wang, US-20200184200, discloses "A biometric classification system includes a
biometric capture system that captures a biometric identifier. A classifier includes a
plurality of group classifiers. Each group classifier in the plurality of classifiers includes a
group discriminator that determines, based on the captured biometric identifier, whether
the biometric identifier belongs to a group of persons associated with the group
discriminator, and includes a plurality of object discriminators. Each object discriminator
is associated with a single person within the group of persons. The group discriminator
determines whether the biometric identifier belongs to the group of persons. The object
discriminator determines whether the biometric identifier belongs to the single person
associated with the object discriminator." (Abstract)
Heven, US-20210358115, discloses "There are provided a system and method of
training a neural network system for anomaly detection, comprising: obtaining a training
dataset including a set of original images and a set of random data vectors; constructing
a neural network system comprising a generator, and a first discriminator and a second
discriminator operatively connected to the generator; training the generator, the first
discriminator and the second discriminator together based on the training dataset, such
that: i) the generator is trained, at least based on evaluation of the first discriminator, to
generate synthetic images meeting a criterion of photo-realism as compared to
corresponding original images; and ii) the second discriminator is trained based on the
original images and the synthetic images to discriminate images with anomaly from
images without anomaly with a given level of accuracy, thereby giving rise to a trained
neural network system." (Abstract)
Kaynig-Fittkau, US-20210158093, discloses "The present disclosure relates to
systems, methods, and non-transitory computer readable media for generating diverse
and realistic synthetic documents using deep learning. In particular, the disclosed
systems can utilize a trained neural network to generate realistic image layouts
comprising page elements that comply with layout parameters. The disclosed systems
can also generate synthetic content corresponding to the page elements within the
image layouts. The disclosed systems insert the synthetic content into the
corresponding page elements of documents based on the image layouts to generate
synthetic documents." (Abstract)
Tran, US-10928830, discloses "Smart car method to navigate a road includes
detecting road-pavement markings using a camera and a sensor; creating a 3D model
based on outputs of the camera and sensor; and navigating the road with a vehicle."
(Abstract)
Sakai, US-20200184269, discloses "A machine learning system includes a first
determination model that determines whether an input image is a second domain
image, and a second determination model that determines whether an extracted image
obtained by extracting an area where an object is presented from the input image is
from the second domain image. Either a pseudo second domain image or the second
domain image is selected and input into the first determination model, and either a first
extracted image in the pseudo second domain image or a second extracted image in
the second domain image is selected and input into an image extracting unit. Learning
of the first determination model is performed based on a first determination result,
learning of the second determination model is performed based on a second
determination result, and learning of a pseudo image generative model is performed
based on the first and second determination results." (Abstract)
Karki, US-20220254022, discloses "A method includes receiving, from a patient,
an image having a visible lesion, modifying the image to appear as if the lesion were not
present, thereby forming a second image, generating a delineation of the abnormality
using a difference between the first and second images, and tagging the segmented
lesions." (Abstract)
Katoh, US-20200242412, discloses "An anomaly detection apparatus performs
training for the generator and the discriminator such that the generator maximizes a
discrimination error of the discriminator and the discriminator minimizes the
discrimination error The anomaly detection apparatus stores, while the training is being
performed, a state of the generator that is half-trained and satisfies a pre-set condition,
and retrains the discriminator by using an image generated by the half-trained generator
that has the stored state." (Abstract)
Katoh, US-20200242399, discloses "An anomaly detection apparatus generates
pieces of image data using a generator and train the generator and a discriminator that
discriminates whether an image data, generated by the generator, is real or fake. The
anomaly detection apparatus trains the generator such that the generator, in generating
the pieces of image data to maximize the discrimination error of the discriminator,
generate at least a piece of specified image data to reduce the discrimination error at a
fixed rate with respect to the pieces of image data and trains, based on the pieces of
image data and the at least a piece of specified image data, the discriminator to
minimize the discrimination error." (Abstract)
Yang, US-20230061517, discloses "This document describes techniques and
apparatuses for verifying the authenticity of images. In aspects, methods include
receiving, by a decoder system (220), an image (210) to be verified; performing feature
recognition on the received image to determine determined features (238) of the
received image; generating a first output (236) defining values representing the
determined features; decoding the received image, by a message decoding neural
network (252), to extract a signature (254) embedded in the received image, the
embedded signature representing recovered features (258) of the received image;
generating a second output (256) defining values representing the recovered features;
providing the first output and the second output to a manipulation detection neural
network (272); and generating, by the manipulation detection neural network, an
estimation of an authenticity of the received image utilizing at least the first output and
the second output." (Abstract)
Lee, US-20210158815, discloses "Disclosed are a method and apparatus for
remotely controlling an imaging apparatus. A method of controlling a remote control
apparatus includes converting a spoken utterance of a user into an utterance text or
receiving the utterance text, applying a generative model-based first learning model to
the utterance text and generating an image having attributes corresponding to a context
of the utterance text, and externally transmitting the image and the utterance text. In
addition, a method of controlling an imaging apparatus includes receiving a first input
including text or speech data and a second input including a first image, capturing at
least one second image based on the first input, comparing the first image and the
second image, and transmitting the second image in response to a comparison result of
the first image and the second image." (Abstract)
Pic, US-20220262150, discloses "An image processing method, for an identity
document that comprises a data page, comprising comprises acquiring a digital image
of the page of data of the identity document. The method further comprises assigning a
class or a super-class to the candidate identity document via automatic classification of
the digital image by a machine-learning algorithm trained beforehand on a set of
reference images in a training phase; processing the digital image to obtain a set of at
least one intermediate image the weight of which is lower than or equal to the weight of
the digital image; applying discrimination to the intermediate image using a discriminator
neural network; and generating an output signal as output from the discriminator neural
network, the value of which is representative of the probability that the candidate identity
document is an authentic document or a fake." (Abstract)
Liu, US-20210358164, discloses "Apparatuses, systems, and techniques to
facilitate application of a style, for which one or more neural networks have not been
trained by a training framework, from one image to content of another image. In at least
one embodiment, a styled output image is generated by one or more neural networks
based on a style contained in a style image and content of a content image where said
one or more neural networks have not been trained by a training framework on said
style." (Abstract)
Kearney, US-11276151, discloses "Dental images are processed according to a
first machine learning model to determine teeth labels. The teeth labels and image are
processed using a second machine learning model to label anatomy. The anatomy
labels, teeth labels, and image are processed using a third machine learning model to
obtain feature measurements, such as pocket depth and clinical attachment level. The
feature measurements, labels, and image may be input to a fourth machine learning
model to obtain a diagnosis for a periodontal condition. Machine learning models may
further be used to reorient, decontaminate, and restore the image prior to processing. A
machine learning model may be trained with images and randomly generated masks in
order to perform inpainting of dental images with missing information." (Abstract)
Zhang, US-20220180622, discloses "A weakly supervised image semantic
segmentation method based on an intra-class discriminator includes: constructing two
levels of intra-class discriminators for each image-level class to determine whether
pixels belonging to the image class belong to a target foreground or a background, and
using weakly supervised data for training; generating a pixel-level image class label
based on the two levels of intra-class discriminators, and generating and outputting a
semantic segmentation result; and further training an image semantic segmentation
module or network by using the label to obtain a final semantic segmentation model for
an unlabeled input image. By means of the new method, intra-class image information
implied in a feature code is fully mined, foreground and background pixels are
accurately distinguished, and performance of a weakly supervised semantic
segmentation model is significantly improved under the condition of only relying on an
image-level annotation." (Abstract)
Lee, US-20220284584, discloses "Methods for training an algorithm to identify
structural anatomical features, for example of a blood vessel, in a non-contrast
computed tomography (NCT) image are described herein. The algorithm may comprise
an image segmentation algorithm, a random forest classifier, or a generative adversarial
network in examples described herein. In one embodiment, a method comprises
receiving a labelled training set for a machine learning image segmentation algorithm.
The labelled training set comprising a plurality of NCT images, each NCT image of the
plurality of NCT images showing a targeted region of a subject, the targeted region
including at least one blood vessel. The labelled training set further comprises a
corresponding plurality of segmentation masks, each segmentation mask labelling at
least one structural feature of a blood vessel in a corresponding NCT image of the
plurality of NCT images. The method further comprises training a machine learning
image segmentation algorithm, using the plurality of NCT images and the corresponding
plurality of segmentation masks, to learn features of the NCT images that correspond to
structural features of the blood vessels labelled in the segmentation masks, and output
a trained image segmentation model. The method further comprises outputting the
trained image segmentation model usable for identifying structural features of a blood
vessel in an NCT image. Further methods are described herein for identifying
anatomical features from an NCT image, and for establishing training sets. Computing
apparatuses and computer readable media are also described herein." (Abstract)
Rajanna, US-20240420316, discloses "Presented herein are systems and
methods for semantic image retrieval. A computing system may identify a first
biomedical image. The computing system may apply an image retrieval model to the
first biomedical image. The image retrieval model may have a convolution block having
a first plurality of parameters to generate a feature map using the first biomedical image.
The first plurality of parameters may be transferred from a preliminary model. The
image retrieval model may have an encoder having a second plurality of parameters to
generate a first hash code for the first biomedical image based on the feature map. The
computing system may select. from the plurality of second biomedical images
corresponding to a plurality of second hash codes, a subset of second biomedical
images using the first hash code. The computing system may provide the subset of
second biomedical images identified using the first biomedical image." (Abstract)
Tran, US-20210157330, discloses "Smart car method to navigate a road includes
detecting one or more objects using a camera and a sensor to delimit boundaries of a
road; creating a 3D model based on outputs of the camera and sensor; and navigating
the road with a vehicle." (Abstract)
Kim, US-20210125320, discloses "An image processing system includes: an
image signal processor including a first neural network, and processing an input image
by using the first neural network so as to generate a post-processed image; and a
discriminator including a second neural network, and receiving a target image and the
post-processed image, and discriminating the target image and the post-processed
image into a real image and a fake image by using the second neural network, wherein
the second neural network is trained to discriminate the target image as a real image
and to discriminate the post-processed image as a fake image, and the first neural
network is trained in such a manner that the post-processed image is discriminated as a
real image by the second neural network." (Abstract)
Baek, US-20210150287, discloses "An image providing apparatus configured to
generate, by using a first artificial intelligence (AI) network, AI metadata including class
information and at least one class map, in which the class information includes at least
one class corresponding to a type of an object among a plurality of predefined objects
included in a first image and the at least one class map indicates a region
corresponding to each class in the first image, generate an encoded image by encoding
the first image, and output the encoded image and the AI metadata through the output
interface." (Abstract)
Banerjee, US-20210201003, discloses "Machine learning is performed using
synthetic data for neural network training using vectors. Facial images are obtained for
a neural network training dataset. Facial elements from the facial images are encoded
into vector representations of the facial elements. A generative adversarial network
(GAN) generator is trained to provide one or more synthetic vectors based on the one or
more vector representations, wherein the one or more synthetic vectors enable
avoidance of discriminator detection in the GAN. The training a GAN further comprises
determining a generator accuracy using the discriminator. The generator accuracy can
enable a classifier, where the classifier comprises a multi-layer perceptron. Additional
synthetic vectors are generated in the GAN, wherein the additional synthetic vectors
avoid discriminator detection. A machine learning neural network is trained using the
additional synthetic vectors. The training a machine learning neural network further
includes using the one or more synthetic vectors." (Abstract)
Tulyakov, US-20210192198, discloses "A landmark detection system can more
accurately detect landmarks in images using a detection scheme that penalizes for
dispersion parameters, such as variance or scale. The landmark detection system can
be trained using both labeled and unlabeled training data in a semi-supervised
approach. The landmark detection system can further implement tracking of an object
across multiple images using landmark data." (Abstract)
Olender, US-20230076868, discloses "A system for completing a medical image
having at least one obscured region includes an input for receiving a first classification
map generated using an acquired optical coherence tomography (OCT) image having at
least one obscured region, the acquired OCT image acquired using an imaging system
and a pre-processing module coupled to the input and configured to create an obscured
region mask. The pre-processing module also generates a second classification map
that has the at least one obscured region filled in. The system also includes a
generative network coupled to the pre-processing module and configured to generate a
synthetic OCT image based on the second classification map and a post-processing
module coupled to the generative network. The post-processing module is configured to
receive the synthetic OCT image and the acquired OCT image and to generate a
completed image based on the synthetic OCT image and the acquired OCT image."
(Abstract)
Liu, US-20220301106, discloses "Provided are a training method and apparatus
for an image processing image processing model, and an image processing method
and apparatus. The training method comprises: acquiring a sample image and a first
reference image, wherein the information quantity and resolution of the sample image
are respectively lower than those of the first reference image; inputting the sample
image into a generative network in an image processing model, and carrying out super resolution
processing and down-sampling processing on the sample image by means of
the generative network, so as to generate and output at least one result image;
determining the total image loss of the at least one result image according to the first
reference image; and adjusting parameters of the generative network according to the
total image loss, so that the total image loss of at least one result image output by the
adjusted generative network meets an image loss condition." (Abstract)
Kearney, US-11189028, discloses "A machine learning model is trained to
predict pixel spacing, distance, and volumetric measurements. Training images are
obtained by inpainting around an original image and scaling the unpainted image to
obtain the training image having a different pixel spacing than the original image. The
machine learning model may include an encoder, a transformer, a first TC layer, and a
second TC layer. During training, loss may be obtained from a comparison of the output
to the first TC layer to a coarse pixel spacing matrix and a comparison of the output of
the second TC layer to a fine pixel spacing matrix. During utilization, the pixel spacing of
an image may be obtained using the machine learning model and used to correct the
image or measurements obtained from the image." (Abstract)
Lee, US-20240242338, discloses "Methods, apparatus and computer readable
media are provided for identifying functional features from a computed tomography (CT)
image. The CT image may be a contrast-enhanced CT image or a non-contrast CT
image. According to some examples, methods, apparatus and computer readable
media are also provided for using machine learning to identify functional features from
CT images. According to some examples, simulated functional image datasets such as
simulated PET images or simulated SUV images are generated from a received CT
image." (Abstract)
Kearney, US-11366985, discloses "In medicine and dentistry, image quality
affects computer vision accuracy. However, some problems are more tolerant of noise
depending on disease severity and radiographic obviousness. There is a need to have a
noise estimation model that adapts to each specific domain. A noise estimation model is
trained to output a set of domain noise estimates for an input image, each estimate
indicating an impact of noise present in the input image on a particular domain, e.g.
labeling of a dental feature such as a dental anatomy, pathology, or treatment. The
noise estimation model is trained by processing image pairs with a set of machine
learning models for a plurality of domains, the image pairs including a raw image and a
modified image obtained by adding noise to the raw image. Outputs of the set of
machine learning models for the raw and modified images are compared to obtain
measured noise metrics. The noise estimation model processes the modified image and
is trained to estimate noise metrics. The noise estimation model is modified according to
differences between the measured noise metrics and estimated noise metrics."
(Abstract)
Mishra, US-20220067519, discloses "Disclosed techniques include neural
network architecture using encoder-decoder models. A facial image is obtained for
processing on a neural network. The facial image includes unpaired facial image
attributes. The facial image is processed through a first encoder-decoder pair and a
second encoder-decoder pair. The first encoder-decoder pair decomposes a first image
attribute subspace. The second encoder-decoder pair decomposes a second image
attribute subspace. The first encoder-decoder pair outputs a transformation mask based
on the first image attribute subspace. The second encoder-decoder pair outputs a
second image transformation mask based on the second image attribute subspace. The
first image transformation mask and the second image transformation mask are
concatenated to enable downstream processing. The concatenated transformation
masks are processed on a third encoder-decoder pair and a resulting image is output.
The resulting image eliminates a paired training data requirement." (Abstract)
Chai, US-20220101104, discloses "aspects of the present disclosure involve a
system comprising a computer-readable storage medium storing a program and method
for video synthesis. The program and method provide for accessing a primary
generative adversarial network (GAN) comprising a pre-trained image generator, a
motion generator comprising a plurality of neural networks, and a video discriminator;
generating an updated GAN based on the primary GAN, by performing operations
comprising identifying input data of the updated GAN, the input data comprising an
initial latent code and a motion domain dataset, training the motion generator based on
the input data, and adjusting weights of the plurality of neural networks of the primary
GAN based on an output of the video discriminator; and generating a synthesized video
based on the primary GAN and the input data." (Abstract)
Kulkarni, US-20230169632, discloses "Certain aspects and features of this
disclosure relate to semantically-aware image extrapolation. In one example, an input
image is segmented to produce an input segmentation map of object instances in the
input image. An object generation network is used to generate an extrapolated semantic
label map for an extrapolated image. The extrapolated semantic label map includes
instances in the original image and instances that will appear in an outpointed region of
the extrapolated image. A panoptic label map is derived from coordinates of output
instances in the extrapolated image and used to identify partial instances and
boundaries. Instance-aware context normalization is used to apply one or more
characteristics from the input image to the out painted region to maintain semantic
continuity. The extrapolated image includes the original image and the outpointed
region and can be rendered or stored for future use." (Abstract)
Contact Information
Any inquiry concerning this communication or earlier communications from the examiner should be directed to NANCY BITAR whose telephone number is (571)270-1041. The examiner can normally be reached Mon-Friday from 8:00 am to 5:00 p.m..
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Ms. Jennifer Mehmood can be reached at 571-272-2976. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
NANCY . BITAR
Examiner
Art Unit 2664
/NANCY BITAR/Primary Examiner, Art Unit 2664