DETAILED ACTION
Applicant's amendment of July 8, 2026 overcomes the following:
Specification objections
Claims 1-2 and 16 objections
Rejection of claim 5 under 35 U.S.C. 112(b)
Rejection of claims 27-29 under 35 U.S.C. 101
Applicant has amended claims 1-2, 5, 16, 19, and 24-29. Claims 3, 17, and 23 have been canceled. Claims 30-32 are new. Claims 1-2, 4-16, 18-22, and 24-32 are pending.
Response to Arguments
Applicant’s arguments filed on July 8, 2026 with respect to pending claims have been considered but are moot in view of the new ground(s) of rejection. The amended claims resulted in changes to the scope and contents; therefore, the grounds of rejection are modified accordingly. It is noted that previously applied prior arts remain in effect.
Applicant asserts that “Applicant believes that the Examiner intended to reject claim 23, not claim 22, under 35 U.S.C. § 112(b) as the rejection is applied to claim 23 on page 6 of the Office Action” (Remarks, Pg. 11).
Examiner agrees.
Examiner intended to reject claim 23 under 35 U.S.C. 112(b), not claim 22, as previously indicated on Pg. 6 of the Non- Final Office Action (OA) of April 16, 2026.
Claim Objections
Claim 30-32 are objected to because of the following informalities:
Claim 30 recites the limitation “imaging system according to claims 1” in line 1 of the claim. However, it should recite “imaging system according to claim 1” instead.
Therefore, based on above, for examination purposes the claimed “imaging system according to claims 1” recited in line 1 of the claim will be interpreted as “imaging system according to claim 1”.
Claim 30 further recites the limitation “the information related to a network structure” in line 3 of the claim. However, it is not clear if the claimed “network structure” recited in line 3 of claim 30 is related to the claimed “neural network” previously recited in claim 1, or not, for example.
Therefore, based on above, for examination purposes the claimed “the information related to a network structure” recited in line 3 of claim 30 will be interpreted as “the information related to the network structure”.
Claim 31 recites the limitation “imaging system according to claims 16” in line 1 of the claim. However, it should recite “imaging system according to claim 16” instead.
Therefore, based on above, for examination purposes the claimed “imaging system according to claims 16” recited in line 1 of the claim will be interpreted as “imaging system according to claim 16”.
Claim 31 further recites the limitation “the information related to a network structure” in line 3 of the claim. However, it is not clear if the claimed “network structure” recited in line 3 of claim 31 is related to the claimed “neural network” previously recited in claim 16, or not, for example.
Therefore, based on above, for examination purposes the claimed “the information related to a network structure” recited in line 3 of claim 31 will be interpreted as “the information related to the network structure”.
Claim 32 recites the limitation “imaging system according to any one of claims 19” in line 1 of the claim. However, it should recite “information processing server according to claim 19” instead.
Therefore, based on above, for examination purposes the claimed “imaging system according to any one of claims 19” recited in line 1 of the claim will be interpreted as “information processing server according to claim 19”.
Claim 32 further recites the limitation “the information related to a network structure” in lines 1-2 of the claim. However, it is not clear if the claimed “network structure” recited in lines 1-2 of claim 32 is related to the claimed “neural network” previously recited in claim 19, or not, for example.
Therefore, based on above, for examination purposes the claimed “the information related to a network structure” recited in lines 1-2 of claim 32 will be interpreted as “the information related to the network structure”.
Appropriate correction is required.
Claim Rejections - 35 USC § 102
The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action:
A person shall be entitled to a patent unless –
(a)(1) the claimed invention was patented, described in a printed publication, or in public use, on sale or otherwise available to the public before the effective filing date of the claimed invention.
Claims 1-2, 4-9, 12, 15-16, 18-22, and 24-29 are rejected under 35 U.S.C. 102(a)(1) as being anticipated by Kashu et al. (US PG Publication No. 2021/0256306 A1), hereafter referred to as Kashu.
Regarding claim 1, Kashu discloses an imaging system that performs object detection (Par. [0006]: an apparatus comprising: a storage device configured to store a plurality of dictionary data for respectively detecting a plurality of different objects from an image) on the basis of a neural network (Par. [0033-35]: detector 213 performs processing for detecting a specific area (e.g., person's face or pupil, dog's face or pupil) from an image. There are times when no specific areas are detected and times when a plurality of specific areas are detected. As for the detection technique, any known method such as AdaBoost or a convolutional neural network are be used… a detector that performs object detection by CNN (convolutional neural network) is used for the detector 213, and dictionary data that is used by the detector in detecting each object is a learned parameter generated in advance using machine learning of a CNN), the imaging system comprising:
at least one processor or circuit configured to function as (Par. [0006]: an apparatus comprising: a storage device configured to store a plurality of dictionary data for respectively detecting a plurality of different objects from an image; and at least one processor configured to function as):
training data input unit configured to input training data for the object detection (Par. [0035-36]: detector that performs object detection by CNN (convolutional neural network) is used for the detector 213, and dictionary data that is used by the detector in detecting each object is a learned parameter generated in advance using machine learning of a CNN… Machine learning of a CNN can be performed by any technique. For example, a predetermined computer such as a server may perform machine learning of a CNN, and the image capturing apparatus 100 may acquire the learned CNN from the predetermined computer… For example, learning of a CNN by an object detection unit may be performed, by a predetermined computer performing supervised learning with image data for use in learning as an input… Also, learning of a CNN by a dictionary estimation unit may be performed, by a predetermined computer performing supervised learning with image data for use in learning as an input and with dictionary data corresponding to an object in the image data for use in learning as supervised data. Learned parameters of a CNN are generated in the manner described above. Learning of a CNN may be performed by the image capturing apparatus 100 or the image processing apparatus described above; input training data for the object detection (e.g. apparatus for respectively detecting a plurality of different objects from an image (i.e. an imaging system that performs object detection) includes learning (i.e. training) of a neural network, including a convolutional neural network (CNN), for example, by using image data (i.e. training data) in learning as an input (i.e. input training data for the object detection), as indicated above), for example);
network structure designation unit configured to designate information related to a network structure of the neural network in the object detection (Par. [0033-36]: detector 213 performs processing for detecting a specific area (e.g., person's face or pupil, dog's face or pupil) from an image… As for the detection technique, any known method such as AdaBoost or a convolutional neural network are be used… detector that performs object detection by CNN (convolutional neural network) is used for the detector 213, and dictionary data that is used by the detector in detecting each object is a learned parameter generated in advance using machine learning of a CNN… Machine learning of a CNN can be performed by any technique. For example, a predetermined computer such as a server may perform machine learning of a CNN, and the image capturing apparatus 100 may acquire the learned CNN from the predetermined computer… For example, learning of a CNN by an object detection unit may be performed, by a predetermined computer performing supervised learning with image data for use in learning as an input… Also, learning of a CNN by a dictionary estimation unit may be performed, by a predetermined computer performing supervised learning with image data for use in learning as an input and with dictionary data corresponding to an object in the image data for use in learning as supervised data. Learned parameters of a CNN are generated in the manner described above. Learning of a CNN may be performed by the image capturing apparatus 100 or the image processing apparatus described above; designate information related to a network structure of the neural network in the object detection (e.g. apparatus for respectively detecting a plurality of different objects from an image (i.e. imaging system that performs object detection) includes learning (i.e. training) of a neural network, including a convolutional neural network (CNN), for example, and implementing a detector that performs object detection by using the convolutional neural network (i.e. a network structure of the neural network in the object detection), including parameters (i.e. information, data, values, etc., related to the network structure) learned (i.e. designated, generated, created, computed, calculated, etc.) by using machine learning of the CNN (i.e. designate information related to a network structure of the neural network in the object detection), as indicated above), for example);
dictionary generation unit configured to generate dictionary data for the object detection on the basis of the training data and the information related to the network structure (Par. [0003]: detection of specific objects is performed by inputting images to a detector together with dictionary data obtained by learning the object to be detected. By changing the dictionary data input to the detector, it is possible to detect different types of objects from within an image; Par. [0033-64]: detector 213 performs processing for detecting a specific area (e.g., person's face or pupil, dog's face or pupil) from an image… As for the detection technique, any known method such as AdaBoost or a convolutional neural network are be used… Dictionary data is data in which features of corresponding objects are registered, for example, and control commands for the logic circuits are defined for every object type… dictionary data for every object is stored by a ROM within the computational unit 102 (a plurality of types of dictionary data is stored). Because dictionary data exists for every type of object and specific area, objects of different types can be detected, by switching dictionary data… a detector that performs object detection by CNN (convolutional neural network) is used for the detector 213, and dictionary data that is used by the detector in detecting each object is a learned parameter generated in advance using machine learning of a CNN… For example, learning of a CNN by an object detection unit may be performed, by a predetermined computer performing supervised learning with image data for use in learning as an input and with the position of an object corresponding to the image data for use in learning as supervised data. Also, learning of a CNN by a dictionary estimation unit may be performed, by a predetermined computer performing supervised learning with image data for use in learning as an input and with dictionary data corresponding to an object in the image data for use in learning as supervised data. Learned parameters of a CNN are generated in the manner described above… detection of a plurality of types of objects that is fewer than the total number is applied by switching the dictionary in accordance with settings determined in advance in order to save on processing speed, bus band… the method for controlling dictionary data switching is determined, based on the calculated priorities of the respective dictionary data… the image data and dictionary data generated in step S301 are input to the detector 213, and a specific area of the object corresponding to dictionary data is detected. The dictionary data that is input depends on the dictionary data switching control method determined… In step S403, it is determined whether local dictionary data is defined for the main object. Depending on the object, an entire area or a local area that is part of the entire area is defined as the detection area… the method of setting dictionary data to be used by the detector 213 in object detection with respect to a plurality of frames… The detector 213 refers to dictionary data in order of the tables stored in the RAM, and performs detection of objects corresponding to the respective dictionary data. A table indicating the input order of dictionary data extracted to the RAM is rewritten as required according to the detection state of objects; generate dictionary data for the object detection on the basis of the training data and the information related to the network structure (e.g. apparatus for respectively detecting a plurality of different objects from an image (i.e. imaging system that performs object detection) includes learning (i.e. training) of a neural network, including a convolutional neural network (CNN), for example, and implementing a detector that performs object detection by using the convolutional neural network (i.e. a network structure of the neural network in the object detection), including setting (i.e. generating, creating, computing, calculating, etc.) dictionary data to be used by the detector in the object detection of each object (i.e. generate dictionary data for the object detection) with respect to a plurality of frames, for example, including parameters (i.e. information, data, values, etc., related to the network structure) learned using machine learning (i.e. training data) of the CNN (i.e. generate dictionary data for the object detection on the basis of the training data and the information related to the network structure), as indicated above), for example); and
an imaging device configured to perform the object detection on the basis of the dictionary data generated by the dictionary generation unit and performs predetermined imaging control on an object detected through the object detection (Par. [0003-7]: detection of specific objects is performed by inputting images to a detector together with dictionary data obtained by learning the object to be detected. By changing the dictionary data input to the detector, it is possible to detect different types of objects from within an image… an apparatus comprising: a storage device configured to store a plurality of dictionary data for respectively detecting a plurality of different objects from an image; and at least one processor configured to function as: a detection unit configured to use partial dictionary data of the plurality of dictionary data to detect an object corresponding to the partial dictionary data, with respect to each frame of an image of a plurality of frames obtained by an image capturing device… an apparatus comprising: an image capturing device configured to capture an image; and an image processing apparatus including: a storage device configured to store a plurality of dictionary data for respectively detecting a plurality of different objects from an image; and at least one processor configured to function as: a detection unit configured to use partial dictionary data of the plurality of dictionary data to detect an object corresponding to the partial dictionary data, with respect to each frame of an image of a plurality of frames obtained by an image capturing device; and a switching unit configured to switch the dictionary data to be used by the detection unit in the plurality of frames, according to a result of the object detection by the detection unit; Par. [0029-39]: FIG. 1 is a sectional side view showing the configuration of a digital single-lens reflex camera (hereinafter, also simply camera) 100 serving as a first embodiment of an image capturing apparatus of the disclosure. FIG. 2 is a block diagram showing the electrical configuration of the camera 100 in FIG. 1 … the configuration related to control will be described using FIG. 2. The computational unit 102 is provided with a dedicated circuit for executing specific computational processing at high speed, in addition to a RAM, a ROM and a multi-core CPU that is able to perform parallel processing of multiple tasks. Due to such hardware, the computational unit 102 constitutes a control unit 201, a main object computational unit 202, a tracking computational unit 203, a focus computational unit 204, and an exposure computational unit 205. The control unit 201 controls the various parts of the camera body 101 and the lens unit 120… a detector that performs object detection by CNN (convolutional neural network) is used for the detector 213, and dictionary data that is used by the detector in detecting each object is a learned parameter generated in advance using machine learning of a CNN in an external device (PC) or the image capturing apparatus 100… For example, learning of a CNN by an object detection unit may be performed, by a predetermined computer performing supervised learning with image data for use in learning as an input and with the position of an object corresponding to the image data for use in learning as supervised data. Also, learning of a CNN by a dictionary estimation unit may be performed, by a predetermined computer performing supervised learning with image data for use in learning as an input and with dictionary data corresponding to an object in the image data for use in learning as supervised data. Learned parameters of a CNN are generated in the manner described above. Learning of a CNN may be performed by the image capturing apparatus 100 or the image processing apparatus described above… tracking computational unit 203 tracks the main object area, based on detection information of the main object. The focus computational unit 204 calculates control values of the focus lens 121 for focusing on the main object area. Also, the exposure computational unit 205 calculates control values of the diaphragm 122 and the image sensor 104 for correctly exposing the main object area… An operation unit 106 is provided with a shutter release switch, a mode dial and the like, and the control unit 201 is able to receive shooting instructions, mode change instructions and other such instructions from the user through the operation unit 106; Par. [0050-69]: the image data and dictionary data generated in step S301 are input to the detector 213, and a specific area of the object corresponding to dictionary data is detected. The dictionary data that is input depends on the dictionary data switching control method determined in step S309… In step S403, it is determined whether local dictionary data is defined for the main object. Depending on the object, an entire area or a local area that is part of the entire area is defined as the detection area… In step S504, the control schedule of dictionary data is determined, based on the detection cycles of respective dictionary data determined in the preceding steps and the detection frequency per frame… a specific example of control of dictionary data will be described using FIGS. 6A and 6B and FIG. 7. Types of dictionary data, detection areas, and processing restrictions on the detector in a control case example envisaged in the present embodiment are shown in FIG. 6A… six types are prepared as dictionary data, in order to detect objects… The priorities and detection cycles of respective dictionary data are determined by both steps S308 and S309 of FIG. 4 and FIG. 5, on the basis of the conditions in FIG. 6A. The results are shown in FIG. 6B. FIG. 6B shows the detection cycles of respective dictionaries determined according to the main object that is detected; and an imaging device configured to perform the object detection on the basis of the dictionary data generated by the dictionary generation unit and performs predetermined imaging control on an object detected through the object detection (e.g. apparatus for respectively detecting a plurality of different objects from an image (i.e. imaging system that performs object detection) includes learning (i.e. training) of a neural network, including a convolutional neural network, for example, and implementing a detector that performs object detection by using the convolutional neural network, including setting (i.e. generating, creating, computing, calculating, etc.) dictionary data to be used by the detector in object detection with respect to each frame of an image of a plurality of frames obtained by an image capturing device, including a camera (i.e. an imaging device configured to perform the object detection on the basis of the dictionary data generated), for example, including dictionary data that is used by the object detector in detecting each object as a learned parameter (i.e. information, data, values, etc., related to the network structure) generated using machine learning of the CNN , for example, including performing control of dictionary data and control commands for logic circuits defined for every object type, for example, including controls of various parts of the image capturing device or camera (i.e. and performs predetermined imaging control on an object detected through the object detection), as indicated above), for example),
wherein the information related to the network structure includes information related to at least one of an image size of input data, a number of channels of the input data, a number of parameters of the neural network, a memory capacity, a type of a layer and a type of an activation function, and a product-sum operation specification (Par. [0035-36]: a detector that performs object detection by CNN (convolutional neural network) is used for the detector 213, and dictionary data that is used by the detector in detecting each object is a learned parameter generated in advance using machine learning of a CNN in an external device (PC) or the image capturing apparatus 100… Learned parameters of a CNN are generated in the manner described above. Learning of a CNN may be performed by the image capturing apparatus 100 or the image processing apparatus described above; wherein the information related to the network structure includes information related to at least one of an image size of input data, a number of channels of the input data, a number of parameters of the neural network, a memory capacity, a type of a layer and a type of an activation function, and a product-sum operation specification (e.g. apparatus for respectively detecting a plurality of different objects from an image (i.e. imaging system that performs object detection) includes learning (i.e. training) of a neural network (i.e. the network structure), including a convolutional neural network (CNN), for example, and implementing a detector that performs object detection by using the convolutional neural network, including parameters (i.e. a number of parameters of the neural network) learned by using machine learning of the CNN (i.e. wherein the information related to the network structure includes information related to at least one of a number of parameters of the neural network), and performs object detection by using the convolutional neural network as indicated above), for example).
Regarding claim 2, claim 1 is incorporated and Kashu discloses the imaging system (Par. [0006]), wherein the imaging device includes a communication unit configured to receive the dictionary data and performs the object detection on the basis of the dictionary data received by the communication unit (Par. [0006]: an apparatus comprising: a storage device configured to store a plurality of dictionary data for respectively detecting a plurality of different objects from an image; and at least one processor configured to function as: a detection unit configured to use partial dictionary data of the plurality of dictionary data to detect an object corresponding to the partial dictionary data, with respect to each frame of an image of a plurality of frames obtained by an image capturing device; Par. [0033-34]: detector 213 performs processing for detecting a specific area (e.g., person's face or pupil, dog's face or pupil) from an image… as for the mode of implementing the method, a program that runs on a CPU, dedicated hardware, or a combination thereof may be used. Also, the type of detection object can be changed by switching the dictionary data that is input to the detector 213… For example, the dictionary data priority calculation unit 211 calculates the priority of dictionary data every predetermined period such as every frame, and, based on the computation result, the dictionary data switching control unit 212 determines dictionary data to be input to the detector 213; Par. [0035-36]: a detector that performs object detection by CNN (convolutional neural network) is used for the detector 213, and dictionary data that is used by the detector in detecting each object is a learned parameter generated in advance using machine learning of a CNN in an external device (PC) or the image capturing apparatus 100… Machine learning of a CNN can be performed by any technique. For example, a predetermined computer such as a server may perform machine learning of a CNN, and the image capturing apparatus 100 may acquire the learned CNN from the predetermined computer; Par. [0050-72]: the image data and dictionary data generated in step S301 are input to the detector 213, and a specific area of the object corresponding to dictionary data is detected. The dictionary data that is input depends on the dictionary data switching control method determined in step S309… the dictionary data corresponding to the main object that is detected is input to the detector 213 in all the frames (per frame from this time) from the sixth frame onward, stable object tracking becomes possible. Dictionary data other than the main object is also input to the detector, and thus even when types of objects other than the main object appears within the shot image, these objects are detectable, and it also becomes possible to change the main object to another type of object… Corresponding local dictionary data is also input to the detector at a detection cycle; receive the dictionary data and performs the object detection on the basis of the dictionary data received (e.g. apparatus for respectively detecting a plurality of different objects from an image (i.e. imaging system that performs object detection) includes learning (i.e. training) of a neural network, including a convolutional neural network, for example, and implementing a detector that performs object detection by using the convolutional neural network, including setting (i.e. generating, creating, computing, calculating, etc.) dictionary data to be used by the detector in object detection with respect to each frame of an image of a plurality of frames obtained (i.e. received, acquired, etc.) by an image capturing device, such as dictionary data to be input to the detector, for example, and a storage device configured to store a plurality of dictionary data for respectively detecting a plurality of different objects from an image (i.e. receive the dictionary data and performs the object detection on the basis of the dictionary data received), as indicated above), for example).
Regarding claim 4, claim 1 is incorporated and Kashu discloses the imaging system (Par. [0006]), wherein the dictionary generation unit is included in an information processing server that is different from the imaging device (Par. [0035-36]: detector that performs object detection by CNN (convolutional neural network) is used for the detector 213, and dictionary data that is used by the detector in detecting each object is a learned parameter generated in advance using machine learning of a CNN… Machine learning of a CNN can be performed by any technique. For example, a predetermined computer such as a server may perform machine learning of a CNN, and the image capturing apparatus 100 may acquire the learned CNN from the predetermined computer… For example, learning of a CNN by an object detection unit may be performed, by a predetermined computer performing supervised learning with image data for use in learning as an input… Also, learning of a CNN by a dictionary estimation unit may be performed, by a predetermined computer performing supervised learning with image data for use in learning as an input and with dictionary data corresponding to an object in the image data for use in learning as supervised data. Learned parameters of a CNN are generated in the manner described above. Learning of a CNN may be performed by the image capturing apparatus 100 or the image processing apparatus described above).
Regarding claim 5, claim 4 is incorporated and Kashu discloses the imaging system (Par. [0006]), wherein the information processing server includes at least one processor or circuit configured to function (Par. [0035-36]: a detector that performs object detection by CNN (convolutional neural network) is used for the detector 213, and dictionary data that is used by the detector in detecting each object is a learned parameter generated in advance using machine learning of a CNN in an external device (PC) or the image capturing apparatus 100… Machine learning of a CNN can be performed by any technique. For example, a predetermined computer such as a server may perform machine learning of a CNN, and the image capturing apparatus 100 may acquire the learned CNN from the predetermined computer… Learning of a CNN may be performed by the image capturing apparatus 100 or the image processing apparatus described above) as Par. [0006]: an apparatus comprising: a storage device configured to store a plurality of dictionary data for respectively detecting a plurality of different objects from an image; and at least one processor configured to function as):
training data acquisition unit configured to acquire the training data for the object detection (Par. [0035-36]: detector that performs object detection by CNN (convolutional neural network) is used for the detector 213, and dictionary data that is used by the detector in detecting each object is a learned parameter generated in advance using machine learning of a CNN… Machine learning of a CNN can be performed by any technique. For example, a predetermined computer such as a server may perform machine learning of a CNN, and the image capturing apparatus 100 may acquire the learned CNN from the predetermined computer… For example, learning of a CNN by an object detection unit may be performed, by a predetermined computer performing supervised learning with image data for use in learning as an input… Also, learning of a CNN by a dictionary estimation unit may be performed, by a predetermined computer performing supervised learning with image data for use in learning as an input and with dictionary data corresponding to an object in the image data for use in learning as supervised data. Learned parameters of a CNN are generated in the manner described above. Learning of a CNN may be performed by the image capturing apparatus 100 or the image processing apparatus described above; acquire the training data for the object detection (e.g. apparatus for respectively detecting a plurality of different objects from an image (i.e. an imaging system that performs object detection) includes learning (i.e. training) of a neural network, including a convolutional neural network (CNN), for example, by using image data (i.e. training data for the object detection) in learning as an input (i.e. acquire the training data for the object detection), as indicated above), for example),
network structure acquisition unit configured to acquire the information related to the network structure (Par. [0033-36]: detector 213 performs processing for detecting a specific area (e.g., person's face or pupil, dog's face or pupil) from an image… As for the detection technique, any known method such as AdaBoost or a convolutional neural network are be used… detector that performs object detection by CNN (convolutional neural network) is used for the detector 213, and dictionary data that is used by the detector in detecting each object is a learned parameter generated in advance using machine learning of a CNN… Machine learning of a CNN can be performed by any technique. For example, a predetermined computer such as a server may perform machine learning of a CNN, and the image capturing apparatus 100 may acquire the learned CNN from the predetermined computer… For example, learning of a CNN by an object detection unit may be performed, by a predetermined computer performing supervised learning with image data for use in learning as an input… Also, learning of a CNN by a dictionary estimation unit may be performed, by a predetermined computer performing supervised learning with image data for use in learning as an input and with dictionary data corresponding to an object in the image data for use in learning as supervised data. Learned parameters of a CNN are generated in the manner described above. Learning of a CNN may be performed by the image capturing apparatus 100 or the image processing apparatus described above; acquire the information related to the network structure (e.g. apparatus for respectively detecting a plurality of different objects from an image (i.e. imaging system that performs object detection) includes learning (i.e. training) of a neural network, including a convolutional neural network (CNN), for example, and implementing a detector that performs object detection by using the convolutional neural network (i.e. a network structure of the neural network in the object detection), including parameters (i.e. information, data, values, etc., related to the network structure) learned (i.e. designated, generated, created, computed, calculated, etc.) by using machine learning of the CNN (i.e. acquire the information related to the network structure), as indicated above), for example),
the dictionary generation unit (Par. [0033-64]: detector 213 performs processing for detecting a specific area (e.g., person's face or pupil, dog's face or pupil) from an image… As for the detection technique, any known method such as AdaBoost or a convolutional neural network are be used… Dictionary data is data in which features of corresponding objects are registered, for example, and control commands for the logic circuits are defined for every object type… dictionary data for every object is stored by a ROM within the computational unit 102 (a plurality of types of dictionary data is stored). Because dictionary data exists for every type of object and specific area, objects of different types can be detected, by switching dictionary data… a detector that performs object detection by CNN (convolutional neural network) is used for the detector 213, and dictionary data that is used by the detector in detecting each object is a learned parameter generated in advance using machine learning of a CNN… For example, learning of a CNN by an object detection unit may be performed, by a predetermined computer performing supervised learning with image data for use in learning as an input and with the position of an object corresponding to the image data for use in learning as supervised data. Also, learning of a CNN by a dictionary estimation unit may be performed, by a predetermined computer performing supervised learning with image data for use in learning as an input and with dictionary data corresponding to an object in the image data for use in learning as supervised data. Learned parameters of a CNN are generated in the manner described above… detection of a plurality of types of objects that is fewer than the total number is applied by switching the dictionary in accordance with settings determined in advance in order to save on processing speed, bus band… the method for controlling dictionary data switching is determined, based on the calculated priorities of the respective dictionary data… the image data and dictionary data generated in step S301 are input to the detector 213, and a specific area of the object corresponding to dictionary data is detected. The dictionary data that is input depends on the dictionary data switching control method determined… In step S403, it is determined whether local dictionary data is defined for the main object. Depending on the object, an entire area or a local area that is part of the entire area is defined as the detection area… the method of setting dictionary data to be used by the detector 213 in object detection with respect to a plurality of frames… The detector 213 refers to dictionary data in order of the tables stored in the RAM, and performs detection of objects corresponding to the respective dictionary data. A table indicating the input order of dictionary data extracted to the RAM is rewritten as required according to the detection state of objects; the dictionary generation (e.g. apparatus for respectively detecting a plurality of different objects from an image (i.e. imaging system that performs object detection) includes learning (i.e. training) of a neural network, including a convolutional neural network (CNN), for example, and implementing a detector that performs object detection by using the convolutional neural network (i.e. a network structure of the neural network in the object detection), including setting (i.e. generating, creating, computing, calculating, etc.) dictionary data to be used by the detector in the object detection of each object (i.e. the dictionary generation) with respect to a plurality of frames, for example, including parameters (i.e. information, data, values, etc., related to the network structure) learned using machine learning (i.e. training data) of the CNN, as indicated above), for example), and
dictionary data transmission unit configured to transmit the dictionary data generated by the dictionary generation unit to the imaging device (Par. [0006-7]: apparatus comprising: a storage device configured to store a plurality of dictionary data for respectively detecting a plurality of different objects from an image; and at least one processor configured to function as: a detection unit configured to use partial dictionary data of the plurality of dictionary data to detect an object corresponding to the partial dictionary data, with respect to each frame of an image of a plurality of frames obtained by an image capturing device; and a switching unit configured to switch the dictionary data to be used by the detection unit in the plurality of frames, according to a result of the object detection… apparatus comprising: an image capturing device configured to capture an image; and an image processing apparatus including: a storage device configured to store a plurality of dictionary data for respectively detecting a plurality of different objects from an image; and at least one processor configured to function as: a detection unit configured to use partial dictionary data of the plurality of dictionary data to detect an object corresponding to the partial dictionary data, with respect to each frame of an image of a plurality of frames obtained by an image capturing device; and a switching unit configured to switch the dictionary data to be used by the detection unit in the plurality of frames, according to a result of the object detection; Par. [033-36]: a program that runs on a CPU, dedicated hardware, or a combination thereof may be used. Also, the type of detection object can be changed by switching the dictionary data that is input to the detector 213… Dictionary data is data in which features of corresponding objects are registered, for example, and control commands for the logic circuits are defined for every object type… dictionary data for every object is stored by a ROM within the computational unit 102 (a plurality of types of dictionary data is stored). Because dictionary data exists for every type of object and specific area, objects of different types can be detected, by switching dictionary data. For example, the dictionary data priority calculation unit 211 calculates the priority of dictionary data every predetermined period such as every frame, and, based on the computation result, the dictionary data switching control unit 212 determines dictionary data to be input to the detector 213… a detector that performs object detection by CNN (convolutional neural network) is used for the detector 213, and dictionary data that is used by the detector in detecting each object is a learned parameter generated in advance using machine learning of a CNN in an external device (PC) or the image capturing apparatus 100… Machine learning of a CNN can be performed by any technique. For example, a predetermined computer such as a server may perform machine learning of a CNN, and the image capturing apparatus 100 may acquire the learned CNN from the predetermined computer; and transmit the dictionary data generated to the imaging device (e.g. apparatus for respectively detecting a plurality of different objects from an image (i.e. imaging system that performs object detection) includes learning (i.e. training) of a neural network, including a convolutional neural network (CNN), for example, and implementing a detector that performs object detection by using the convolutional neural network (i.e. a network structure of the neural network in the object detection), including setting (i.e. generating, creating, computing, calculating, etc.) dictionary data to be used by the detector in the object detection of each object (i.e. the dictionary data generated) with respect to a plurality of frames obtained by an image capturing device, including a camera, for example, including a plurality of types of dictionary data stored by a ROM within a computational unit, such as a storage device configured to store a plurality of dictionary data for respectively detecting a plurality of different objects from an image and at least one processor configured to switch the dictionary data to be used by the detection (i.e. and transmit the dictionary data generated to the imaging device), as indicated above), for example).
Regarding claim 6, claim 1 is incorporated and Kashu discloses the imaging system (Par. [0006]), wherein the dictionary generation unit is configured to select a dictionary suitable for an object of the training data from among a plurality of pieces of the dictionary data prepared in advance (Par. [0035-36]: a detector that performs object detection by CNN (convolutional neural network) is used for the detector 213, and dictionary data that is used by the detector in detecting each object is a learned parameter generated in advance using machine learning of a CNN in an external device (PC) or the image capturing apparatus 100… Machine learning of a CNN can be performed by any technique. For example, a predetermined computer such as a server may perform machine learning of a CNN, and the image capturing apparatus 100 may acquire the learned CNN from the predetermined computer; Par. [0064-65]: a plurality of methods may be stored in advance in the ROM within the computational unit 102, in the form of data tables indicating the combinations of dictionary data to be used by the detector 213 in respective frames. The control unit 201 selects a table according to settings (people priority, animal (dog, cat, bird) priority, vehicle (two-wheeled, four-wheeled) priority, etc.) configured by the user as to which object to detect as object detection, for example, and extracts the selected table to the RAM… If there is such dictionary data, the processing advances to step S502, and the detection cycle of the dictionary thereof is set to 1 [frame/detection] (select as dictionary data for acquiring per frame detection results).
Regarding claim 7, claim 1 is incorporated and Kashu discloses the imaging system (Par. [0006]), wherein the dictionary generation unit is configured to generate the dictionary data by performing learning on the basis of the training data (Par. [0003]: detection of specific objects is performed by inputting images to a detector together with dictionary data obtained by learning the object to be detected. By changing the dictionary data input to the detector, it is possible to detect different types of objects from within an image; Par. [0033-64]: detector 213 performs processing for detecting a specific area (e.g., person's face or pupil, dog's face or pupil) from an image… As for the detection technique, any known method such as AdaBoost or a convolutional neural network are be used… Dictionary data is data in which features of corresponding objects are registered, for example, and control commands for the logic circuits are defined for every object type… dictionary data for every object is stored by a ROM within the computational unit 102 (a plurality of types of dictionary data is stored). Because dictionary data exists for every type of object and specific area, objects of different types can be detected, by switching dictionary data… a detector that performs object detection by CNN (convolutional neural network) is used for the detector 213, and dictionary data that is used by the detector in detecting each object is a learned parameter generated in advance using machine learning of a CNN… For example, learning of a CNN by an object detection unit may be performed, by a predetermined computer performing supervised learning with image data for use in learning as an input and with the position of an object corresponding to the image data for use in learning as supervised data. Also, learning of a CNN by a dictionary estimation unit may be performed, by a predetermined computer performing supervised learning with image data for use in learning as an input and with dictionary data corresponding to an object in the image data for use in learning as supervised data. Learned parameters of a CNN are generated in the manner described above… detection of a plurality of types of objects that is fewer than the total number is applied by switching the dictionary in accordance with settings determined in advance in order to save on processing speed, bus band… the method for controlling dictionary data switching is determined, based on the calculated priorities of the respective dictionary data… the image data and dictionary data generated in step S301 are input to the detector 213, and a specific area of the object corresponding to dictionary data is detected. The dictionary data that is input depends on the dictionary data switching control method determined… In step S403, it is determined whether local dictionary data is defined for the main object. Depending on the object, an entire area or a local area that is part of the entire area is defined as the detection area… the method of setting dictionary data to be used by the detector 213 in object detection with respect to a plurality of frames… The detector 213 refers to dictionary data in order of the tables stored in the RAM, and performs detection of objects corresponding to the respective dictionary data. A table indicating the input order of dictionary data extracted to the RAM is rewritten as required according to the detection state of objects; generate the dictionary data by performing learning on the basis of the training data (e.g. apparatus for respectively detecting a plurality of different objects from an image includes learning (i.e. training) of a neural a neural network, including a convolutional neural network, and implementing a detector that performs object detection by using the convolutional neural network (i.e. a network structure of the neural network in the object detection), for example, including setting dictionary data to be used by the detector in object detection (i.e. generate dictionary data for the object detection) with respect to a plurality of frames, for example, including dictionary data that is used by the object detector in detecting each object as a learned parameter generated in advance using machine learning of the CNN (i.e. generate the dictionary data by performing learning on the basis of the training data), as indicated above), for example).
Regarding claim 8, claim 1 is incorporated and Kashu discloses the imaging system (Par. [0006]), wherein the training data input unit and the network structure designation unit are included in an information processing terminal that is different from the imaging device (Par. [0035-36]: detector that performs object detection by CNN (convolutional neural network) is used for the detector 213, and dictionary data that is used by the detector in detecting each object is a learned parameter generated in advance using machine learning of a CNN… Machine learning of a CNN can be performed by any technique. For example, a predetermined computer such as a server may perform machine learning of a CNN, and the image capturing apparatus 100 may acquire the learned CNN from the predetermined computer… For example, learning of a CNN by an object detection unit may be performed, by a predetermined computer performing supervised learning with image data for use in learning as an input… Also, learning of a CNN by a dictionary estimation unit may be performed, by a predetermined computer performing supervised learning with image data for use in learning as an input and with dictionary data corresponding to an object in the image data for use in learning as supervised data. Learned parameters of a CNN are generated in the manner described above. Learning of a CNN may be performed by the image capturing apparatus 100 or the image processing apparatus described above).
Regarding claim 9, claim 1 is incorporated and Kashu discloses the imaging system (Par. [0006]), wherein the training data includes image data (Par. [0006]: an apparatus comprising: a storage device configured to store a plurality of dictionary data for respectively detecting a plurality of different objects from an image; Par. [0033-36]: detector 213 performs processing for detecting a specific area (e.g., person's face or pupil, dog's face or pupil) from an image… As for the detection technique, any known method such as AdaBoost or a convolutional neural network are be used… detector that performs object detection by CNN (convolutional neural network) is used for the detector 213, and dictionary data that is used by the detector in detecting each object is a learned parameter generated in advance using machine learning of a CNN… Machine learning of a CNN can be performed by any technique. For example, a predetermined computer such as a server may perform machine learning of a CNN, and the image capturing apparatus 100 may acquire the learned CNN from the predetermined computer… For example, learning of a CNN by an object detection unit may be performed, by a predetermined computer performing supervised learning with image data for use in learning as an input… Also, learning of a CNN by a dictionary estimation unit may be performed, by a predetermined computer performing supervised learning with image data for use in learning as an input and with dictionary data corresponding to an object in the image data for use in learning as supervised data. Learned parameters of a CNN are generated in the manner described above. Learning of a CNN may be performed by the image capturing apparatus 100 or the image processing apparatus described above) and region information of the image data where a target object is present (Par. [0054-69]: a method of calculating the priorities of dictionary data in step S308 of FIG. 3 will be described using FIG. 4… In step S403, it is determined whether local dictionary data is defined for the main object. Depending on the object, an entire area or a local area that is part of the entire area is defined as the detection area… the detection cycles of respective dictionary data that were determined in the previous frame are initialized. Here, the detection cycle is a parameter [frame/detection] representing the number of frames per detection result acquisition… a specific example of control of dictionary data will be described using FIGS. 6A and 6B and FIG. 7. Types of dictionary data, detection areas, and processing restrictions on the detector in a control case example envisaged in the present embodiment are shown in FIG. 6A. In the present embodiment, six types are prepared as dictionary data, in order to detect objects… it is possible to process three types of dictionaries per frame… The priorities and detection cycles of respective dictionary data are determined).
Regarding claim 12, claim 1 is incorporated and Kashu discloses the imaging system (Par. [0006]), further comprising:
display unit that displays a result of the object detection as a frame in a superimposed manner on an image from the imaging device (Par. [0089-102]: in a shooting mode in which immediacy is required in display and tracking control during continuous shooting and the like, a continuity check is not performed, and if an object is detected even once, the detection result is immediately employed as the main object, and reflection on tracking control, dictionary data control and frame display is performed… in a shooting mode in which continuity checks are not performed such as continuous shooting mode, the dictionary data of all the types continues to be applied even after an object is detected once and employed as the main object, and when an object is continuously detected a given number of times or more, that object is reemployed as the correct object. Tracking control, dictionary data control and frame display control also change according to the main object that has been reemployed… In step S1401, it is checked whether there is an object detection result within the input image. If there is an object detection result, the processing advances to step S1402… In FIGS. 16A to 16C, even in the case where part of the two-wheeled vehicle is detected as the dog 1505 in the first frame, the same dictionary data control as when there is no main object (the dictionary data of all the classification is set) is performed for a given period from the second frame, rather than performing dictionary data control conforming to dog selected as the main object. In FIGS. 16A to 16C, the detection result of a two-wheeled vehicle 1506, which is the correct detection result, is obtained a plurality of times from the second frame, and the continuous detection frequency exceeds the threshold value (4 in FIGS. 16A to 16C) of the continuous detection frequency in the eighth frame. The main object is then changed from dog to two-wheeled in the ninth frame, and dictionary data control is also changed to conform to two-wheeled. By reemploying the new object detection result detected a plurality of times as the main object, erroneous detection can be corrected while maintaining immediacy during continuous shooting, and the risk of erroneous tracking can be suppressed; display unit that displays a result of the object detection as a frame in a superimposed manner on an image from the imaging device (e.g. apparatus for respectively detecting a plurality of different objects from an image includes shooting mode in which immediacy is required in display and tracking control during continuous shooting and displays a result of the object detection as a frame in a superimposed manner on an image from the imaging device, as shown in Fig. 16C below:
PNG
media_image1.png
494
775
media_image1.png
Greyscale
, for example).
Regarding claim 15, claim 1 is incorporated and Kashu discloses the imaging system (Par. [0006]), wherein the imaging device includes training data generation unit configured to generate the training data (Par. [0035-36]: detector that performs object detection by CNN (convolutional neural network) is used for the detector 213, and dictionary data that is used by the detector in detecting each object is a learned parameter generated in advance using machine learning of a CNN… Machine learning of a CNN can be performed by any technique. For example, a predetermined computer such as a server may perform machine learning of a CNN, and the image capturing apparatus 100 may acquire the learned CNN from the predetermined computer… For example, learning of a CNN by an object detection unit may be performed, by a predetermined computer performing supervised learning with image data for use in learning as an input… Also, learning of a CNN by a dictionary estimation unit may be performed, by a predetermined computer performing supervised learning with image data for use in learning as an input and with dictionary data corresponding to an object in the image data for use in learning as supervised data. Learned parameters of a CNN are generated in the manner described above. Learning of a CNN may be performed by the image capturing apparatus 100 or the image processing apparatus described above; wherein the imaging device includes training data generation unit configured to generate the training data (e.g. apparatus for respectively detecting a plurality of different objects from an image includes learning (i.e. training) of a neural network, including a convolutional neural network (CNN), for example, including setting dictionary data to be used by the detector in object detection with respect to each frame of an image of a plurality of frames obtained by an image capturing device (i.e. the imaging device), for example, by performing learning with image data for use in learning as an input and with dictionary data corresponding to an object in the image data for use in learning as supervised data (i.e. generate the training data), as indicated above), for example).
Regarding claim 16, Kashu discloses an imaging device that performs object detection (Par. [0006]: an apparatus comprising: a storage device configured to store a plurality of dictionary data for respectively detecting a plurality of different objects from an image) on the basis of a neural network (Par. [0033-35]: detector 213 performs processing for detecting a specific area (e.g., person's face or pupil, dog's face or pupil) from an image. There are times when no specific areas are detected and times when a plurality of specific areas are detected. As for the detection technique, any known method such as AdaBoost or a convolutional neural network are be used… a detector that performs object detection by CNN (convolutional neural network) is used for the detector 213, and dictionary data that is used by the detector in detecting each object is a learned parameter generated in advance using machine learning of a CNN), the imaging device comprising:
at least one processor or circuit configured to function as (Par. [0006]: an apparatus comprising: a storage device configured to store a plurality of dictionary data for respectively detecting a plurality of different objects from an image; and at least one processor configured to function as):
training data input unit configured to input training data for the object detection (Par. [0035-36]: detector that performs object detection by CNN (convolutional neural network) is used for the detector 213, and dictionary data that is used by the detector in detecting each object is a learned parameter generated in advance using machine learning of a CNN… Machine learning of a CNN can be performed by any technique. For example, a predetermined computer such as a server may perform machine learning of a CNN, and the image capturing apparatus 100 may acquire the learned CNN from the predetermined computer… For example, learning of a CNN by an object detection unit may be performed, by a predetermined computer performing supervised learning with image data for use in learning as an input… Also, learning of a CNN by a dictionary estimation unit may be performed, by a predetermined computer performing supervised learning with image data for use in learning as an input and with dictionary data corresponding to an object in the image data for use in learning as supervised data. Learned parameters of a CNN are generated in the manner described above. Learning of a CNN may be performed by the image capturing apparatus 100 or the image processing apparatus described above; input training data for the object detection (e.g. apparatus for respectively detecting a plurality of different objects from an image (i.e. an imaging system that performs object detection) includes learning (i.e. training) of a neural network, including a convolutional neural network (CNN), for example, by using image data (i.e. training data) in learning as an input (i.e. input training data for the object detection), as indicated above), for example);
network structure designation unit configured to designate information related to a network structure of the neural network in the object detection (Par. [0033-36]: detector 213 performs processing for detecting a specific area (e.g., person's face or pupil, dog's face or pupil) from an image… As for the detection technique, any known method such as AdaBoost or a convolutional neural network are be used… detector that performs object detection by CNN (convolutional neural network) is used for the detector 213, and dictionary data that is used by the detector in detecting each object is a learned parameter generated in advance using machine learning of a CNN… Machine learning of a CNN can be performed by any technique. For example, a predetermined computer such as a server may perform machine learning of a CNN, and the image capturing apparatus 100 may acquire the learned CNN from the predetermined computer… For example, learning of a CNN by an object detection unit may be performed, by a predetermined computer performing supervised learning with image data for use in learning as an input… Also, learning of a CNN by a dictionary estimation unit may be performed, by a predetermined computer performing supervised learning with image data for use in learning as an input and with dictionary data corresponding to an object in the image data for use in learning as supervised data. Learned parameters of a CNN are generated in the manner described above. Learning of a CNN may be performed by the image capturing apparatus 100 or the image processing apparatus described above; designate information related to a network structure of the neural network in the object detection (e.g. apparatus for respectively detecting a plurality of different objects from an image (i.e. imaging system that performs object detection) includes learning (i.e. training) of a neural network, including a convolutional neural network (CNN), for example, and implementing a detector that performs object detection by using the convolutional neural network (i.e. a network structure of the neural network in the object detection), including parameters (i.e. information, data, values, etc., related to the network structure) learned (i.e. designated, generated, created, computed, calculated, etc.) by using machine learning of the CNN (i.e. designate information related to a network structure of the neural network in the object detection), as indicated above), for example);
communication unit configured to transmit the training data and the information related to the network structure to an information processing server (Par. [0035-36]: a detector that performs object detection by CNN (convolutional neural network) is used for the detector 213, and dictionary data that is used by the detector in detecting each object is a learned parameter generated in advance using machine learning of a CNN in an external device (PC) or the image capturing apparatus 100… Machine learning of a CNN can be performed by any technique. For example, a predetermined computer such as a server may perform machine learning of a CNN, and the image capturing apparatus 100 may acquire the learned CNN from the predetermined computer); and
imaging control unit configured to acquire dictionary data for the object detection generated on the basis of the training data and the information related to the network structure in the information processing server from the information processing server via the communication unit (Par. [0006]: an apparatus comprising: a storage device configured to store a plurality of dictionary data for respectively detecting a plurality of different objects from an image; and at least one processor configured to function as: a detection unit configured to use partial dictionary data of the plurality of dictionary data to detect an object corresponding to the partial dictionary data, with respect to each frame of an image of a plurality of frames obtained by an image capturing device; Par. [0033-34]: detector 213 performs processing for detecting a specific area (e.g., person's face or pupil, dog's face or pupil) from an image… as for the mode of implementing the method, a program that runs on a CPU, dedicated hardware, or a combination thereof may be used. Also, the type of detection object can be changed by switching the dictionary data that is input to the detector 213… For example, the dictionary data priority calculation unit 211 calculates the priority of dictionary data every predetermined period such as every frame, and, based on the computation result, the dictionary data switching control unit 212 determines dictionary data to be input to the detector 213; Par. [0035-36]: a detector that performs object detection by CNN (convolutional neural network) is used for the detector 213, and dictionary data that is used by the detector in detecting each object is a learned parameter generated in advance using machine learning of a CNN in an external device (PC) or the image capturing apparatus 100… Machine learning of a CNN can be performed by any technique. For example, a predetermined computer such as a server may perform machine learning of a CNN, and the image capturing apparatus 100 may acquire the learned CNN from the predetermined computer; Par. [0050-72]: the image data and dictionary data generated in step S301 are input to the detector 213, and a specific area of the object corresponding to dictionary data is detected. The dictionary data that is input depends on the dictionary data switching control method determined in step S309… the dictionary data corresponding to the main object that is detected is input to the detector 213 in all the frames (per frame from this time) from the sixth frame onward, stable object tracking becomes possible. Dictionary data other than the main object is also input to the detector, and thus even when types of objects other than the main object appears within the shot image, these objects are detectable, and it also becomes possible to change the main object to another type of object… Corresponding local dictionary data is also input to the detector at a detection cycle; acquire dictionary data for the object detection generated on the basis of the training data and the information related to the network structure in the information processing server from the information processing server (e.g. apparatus for respectively detecting a plurality of different objects from an image (i.e. imaging system that performs object detection) includes learning (i.e. training) of a neural network, including a convolutional neural network, for example, and implementing a detector that performs object detection by using the convolutional neural network, including setting (i.e. generating, creating, computing, calculating, etc.) dictionary data to be used by the detector in object detection with respect to each frame of an image of a plurality of frames obtained by an image capturing device, such as dictionary data to be input to the detector, for example, including dictionary data that is used by the object detector in detecting each object as a learned parameter (i.e. information, data, values, etc., related to the network structure) generated using machine learning of the CNN (i.e. generated on the basis of the training data and the information related to the network structure), for example, and a storage device configured to store a plurality of dictionary data for respectively detecting a plurality of different objects from an image, for example, including a predetermined computer such as a server to perform machine learning of the CNN, and the image capturing apparatus acquires the learned CNN from the predetermined computer (i.e. acquire dictionary data for the object detection generated on the basis of the training data and the information related to the network structure in the information processing server from the information processing serve), as indicated above), for example), performing the object detection on the basis of the dictionary data, and performing predetermined imaging control on an object detected through the object detection (Par. [0003-7]: detection of specific objects is performed by inputting images to a detector together with dictionary data obtained by learning the object to be detected. By changing the dictionary data input to the detector, it is possible to detect different types of objects from within an image… an apparatus comprising: a storage device configured to store a plurality of dictionary data for respectively detecting a plurality of different objects from an image; and at least one processor configured to function as: a detection unit configured to use partial dictionary data of the plurality of dictionary data to detect an object corresponding to the partial dictionary data, with respect to each frame of an image of a plurality of frames obtained by an image capturing device… an apparatus comprising: an image capturing device configured to capture an image; and an image processing apparatus including: a storage device configured to store a plurality of dictionary data for respectively detecting a plurality of different objects from an image; and at least one processor configured to function as: a detection unit configured to use partial dictionary data of the plurality of dictionary data to detect an object corresponding to the partial dictionary data, with respect to each frame of an image of a plurality of frames obtained by an image capturing device; and a switching unit configured to switch the dictionary data to be used by the detection unit in the plurality of frames, according to a result of the object detection by the detection unit; Par. [0029-39]: FIG. 1 is a sectional side view showing the configuration of a digital single-lens reflex camera (hereinafter, also simply camera) 100 serving as a first embodiment of an image capturing apparatus of the disclosure. FIG. 2 is a block diagram showing the electrical configuration of the camera 100 in FIG. 1 … the configuration related to control will be described using FIG. 2. The computational unit 102 is provided with a dedicated circuit for executing specific computational processing at high speed, in addition to a RAM, a ROM and a multi-core CPU that is able to perform parallel processing of multiple tasks. Due to such hardware, the computational unit 102 constitutes a control unit 201, a main object computational unit 202, a tracking computational unit 203, a focus computational unit 204, and an exposure computational unit 205. The control unit 201 controls the various parts of the camera body 101 and the lens unit 120… a detector that performs object detection by CNN (convolutional neural network) is used for the detector 213, and dictionary data that is used by the detector in detecting each object is a learned parameter generated in advance using machine learning of a CNN in an external device (PC) or the image capturing apparatus 100… For example, learning of a CNN by an object detection unit may be performed, by a predetermined computer performing supervised learning with image data for use in learning as an input and with the position of an object corresponding to the image data for use in learning as supervised data. Also, learning of a CNN by a dictionary estimation unit may be performed, by a predetermined computer performing supervised learning with image data for use in learning as an input and with dictionary data corresponding to an object in the image data for use in learning as supervised data. Learned parameters of a CNN are generated in the manner described above. Learning of a CNN may be performed by the image capturing apparatus 100 or the image processing apparatus described above… tracking computational unit 203 tracks the main object area, based on detection information of the main object. The focus computational unit 204 calculates control values of the focus lens 121 for focusing on the main object area. Also, the exposure computational unit 205 calculates control values of the diaphragm 122 and the image sensor 104 for correctly exposing the main object area… An operation unit 106 is provided with a shutter release switch, a mode dial and the like, and the control unit 201 is able to receive shooting instructions, mode change instructions and other such instructions from the user through the operation unit 106; Par. [0050-69]: the image data and dictionary data generated in step S301 are input to the detector 213, and a specific area of the object corresponding to dictionary data is detected. The dictionary data that is input depends on the dictionary data switching control method determined in step S309… In step S403, it is determined whether local dictionary data is defined for the main object. Depending on the object, an entire area or a local area that is part of the entire area is defined as the detection area… In step S504, the control schedule of dictionary data is determined, based on the detection cycles of respective dictionary data determined in the preceding steps and the detection frequency per frame… a specific example of control of dictionary data will be described using FIGS. 6A and 6B and FIG. 7. Types of dictionary data, detection areas, and processing restrictions on the detector in a control case example envisaged in the present embodiment are shown in FIG. 6A… six types are prepared as dictionary data, in order to detect objects… The priorities and detection cycles of respective dictionary data are determined by both steps S308 and S309 of FIG. 4 and FIG. 5, on the basis of the conditions in FIG. 6A. The results are shown in FIG. 6B. FIG. 6B shows the detection cycles of respective dictionaries determined according to the main object that is detected; performing the object detection on the basis of the dictionary data, and performing predetermined imaging control on an object detected through the object detection (e.g. apparatus for respectively detecting a plurality of different objects from an image (i.e. imaging system that performs object detection) includes learning (i.e. training) of a neural network, including a convolutional neural network, for example, and implementing a detector that performs object detection by using the convolutional neural network, including setting (i.e. generating, creating, computing, calculating, etc.) dictionary data to be used by the detector in object detection with respect to each frame of an image of a plurality of frames obtained by an image capturing device, including a camera (i.e. performing the object detection on the basis of the dictionary data), for example, including dictionary data that is used by the object detector in detecting each object as a learned parameter (i.e. information, data, values, etc., related to the network structure) generated using machine learning of the CNN , for example, including performing control of dictionary data and control commands for logic circuits defined for every object type, for example, including controls of various parts of the image capturing device or camera (i.e. and performing predetermined imaging control on an object detected through the object detection), as indicated above), for example),
wherein the information related to the network structure includes information related to at least one of an image size of input data, a number of channels of the input data, a number of parameters of the neural network, a memory capacity, a type of a layer and a type of an activation function, and a product-sum operation specification (Par. [0035-36]: a detector that performs object detection by CNN (convolutional neural network) is used for the detector 213, and dictionary data that is used by the detector in detecting each object is a learned parameter generated in advance using machine learning of a CNN in an external device (PC) or the image capturing apparatus 100… Learned parameters of a CNN are generated in the manner described above. Learning of a CNN may be performed by the image capturing apparatus 100 or the image processing apparatus described above; wherein the information related to the network structure includes information related to at least one of an image size of input data, a number of channels of the input data, a number of parameters of the neural network, a memory capacity, a type of a layer and a type of an activation function, and a product-sum operation specification (e.g. apparatus for respectively detecting a plurality of different objects from an image (i.e. imaging system that performs object detection) includes learning (i.e. training) of a neural network (i.e. the network structure), including a convolutional neural network (CNN), for example, and implementing a detector that performs object detection by using the convolutional neural network, including parameters (i.e. a number of parameters of the neural network) learned by using machine learning of the CNN (i.e. wherein the information related to the network structure includes information related to at least one of a number of parameters of the neural network), and performs object detection by using the convolutional neural network as indicated above), for example).
Regarding claim 18, claim 16 is incorporated and Kashu discloses the imaging device (Par. [0006]), further comprising:
display unit that displays a result of the object detection as a frame in a superimposed manner on an image (Par. [0089-102]: in a shooting mode in which immediacy is required in display and tracking control during continuous shooting and the like, a continuity check is not performed, and if an object is detected even once, the detection result is immediately employed as the main object, and reflection on tracking control, dictionary data control and frame display is performed… in a shooting mode in which continuity checks are not performed such as continuous shooting mode, the dictionary data of all the types continues to be applied even after an object is detected once and employed as the main object, and when an object is continuously detected a given number of times or more, that object is reemployed as the correct object. Tracking control, dictionary data control and frame display control also change according to the main object that has been reemployed… In step S1401, it is checked whether there is an object detection result within the input image. If there is an object detection result, the processing advances to step S1402… In FIGS. 16A to 16C, even in the case where part of the two-wheeled vehicle is detected as the dog 1505 in the first frame, the same dictionary data control as when there is no main object (the dictionary data of all the classification is set) is performed for a given period from the second frame, rather than performing dictionary data control conforming to dog selected as the main object. In FIGS. 16A to 16C, the detection result of a two-wheeled vehicle 1506, which is the correct detection result, is obtained a plurality of times from the second frame, and the continuous detection frequency exceeds the threshold value (4 in FIGS. 16A to 16C) of the continuous detection frequency in the eighth frame. The main object is then changed from dog to two-wheeled in the ninth frame, and dictionary data control is also changed to conform to two-wheeled. By reemploying the new object detection result detected a plurality of times as the main object, erroneous detection can be corrected while maintaining immediacy during continuous shooting, and the risk of erroneous tracking can be suppressed; display unit that displays a result of the object detection as a frame in a superimposed manner on an image (e.g. apparatus for respectively detecting a plurality of different objects from an image includes shooting mode in which immediacy is required in display and tracking control during continuous shooting and displays a result of the object detection as a frame in a superimposed manner on an image from the imaging device, as shown in Fig. 16C below:
PNG
media_image1.png
494
775
media_image1.png
Greyscale
, for example).
Regarding claim 19, Kashu discloses an information processing server (Par. [0006]: an apparatus comprising: a storage device configured to store a plurality of dictionary data for respectively detecting a plurality of different objects from an image; and at least one processor configured to function as: a detection unit configured to use partial dictionary data of the plurality of dictionary data to detect an object corresponding to the partial dictionary data, with respect to each frame of an image of a plurality of frames obtained by an image capturing; Par. [0035-36]: detector that performs object detection by CNN (convolutional neural network) is used for the detector 213, and dictionary data that is used by the detector in detecting each object is a learned parameter generated in advance using machine learning of a CNN… Machine learning of a CNN can be performed by any technique. For example, a predetermined computer such as a server may perform machine learning of a CNN, and the image capturing apparatus 100 may acquire the learned CNN from the predetermined computer) comprising:
at least one processor or circuit configured to function as (Par. [0006]: an apparatus comprising: a storage device configured to store a plurality of dictionary data for respectively detecting a plurality of different objects from an image; and at least one processor configured to function as):
training data acquisition unit configured to acquire training data for object detection (Par. [0035-36]: detector that performs object detection by CNN (convolutional neural network) is used for the detector 213, and dictionary data that is used by the detector in detecting each object is a learned parameter generated in advance using machine learning of a CNN… Machine learning of a CNN can be performed by any technique. For example, a predetermined computer such as a server may perform machine learning of a CNN, and the image capturing apparatus 100 may acquire the learned CNN from the predetermined computer… For example, learning of a CNN by an object detection unit may be performed, by a predetermined computer performing supervised learning with image data for use in learning as an input… Also, learning of a CNN by a dictionary estimation unit may be performed, by a predetermined computer performing supervised learning with image data for use in learning as an input and with dictionary data corresponding to an object in the image data for use in learning as supervised data. Learned parameters of a CNN are generated in the manner described above. Learning of a CNN may be performed by the image capturing apparatus 100 or the image processing apparatus described above; acquire training data for object detection (e.g. apparatus for respectively detecting a plurality of different objects from an image (i.e. an imaging system that performs object detection) includes learning (i.e. training) of a neural network, including a convolutional neural network (CNN), for example, by using image data (i.e. training data) in learning as an input (i.e. acquire training data for object detection), as indicated above), for example);
network structure acquisition unit configured to acquire information related to a network structure of an imaging device (Par. [0033-36]: detector 213 performs processing for detecting a specific area (e.g., person's face or pupil, dog's face or pupil) from an image… As for the detection technique, any known method such as AdaBoost or a convolutional neural network are be used… detector that performs object detection by CNN (convolutional neural network) is used for the detector 213, and dictionary data that is used by the detector in detecting each object is a learned parameter generated in advance using machine learning of a CNN… Machine learning of a CNN can be performed by any technique. For example, a predetermined computer such as a server may perform machine learning of a CNN, and the image capturing apparatus 100 may acquire the learned CNN from the predetermined computer… For example, learning of a CNN by an object detection unit may be performed, by a predetermined computer performing supervised learning with image data for use in learning as an input… Also, learning of a CNN by a dictionary estimation unit may be performed, by a predetermined computer performing supervised learning with image data for use in learning as an input and with dictionary data corresponding to an object in the image data for use in learning as supervised data. Learned parameters of a CNN are generated in the manner described above. Learning of a CNN may be performed by the image capturing apparatus 100 or the image processing apparatus described above; acquire information related to a network structure of an imaging device (e.g. apparatus for respectively detecting a plurality of different objects from an image (i.e. imaging system that performs object detection) includes learning (i.e. training) of a neural network, including a convolutional neural network (CNN), for example, and implementing a detector that performs object detection by using the convolutional neural network (i.e. a network structure of the neural network in the object detection), including parameters (i.e. information, data, values, etc., related to the network structure) learned (i.e. acquired, designated, generated, created, computed, calculated, etc.) by using machine learning of the CNN (i.e. acquire information related to a network structure of an imaging device), as indicated above), for example);
dictionary generation unit configured to generate dictionary data for the object detection on the basis of the training data and the information related to the network structure (Par. [0003]: detection of specific objects is performed by inputting images to a detector together with dictionary data obtained by learning the object to be detected. By changing the dictionary data input to the detector, it is possible to detect different types of objects from within an image; Par. [0033-64]: detector 213 performs processing for detecting a specific area (e.g., person's face or pupil, dog's face or pupil) from an image… As for the detection technique, any known method such as AdaBoost or a convolutional neural network are be used… Dictionary data is data in which features of corresponding objects are registered, for example, and control commands for the logic circuits are defined for every object type… dictionary data for every object is stored by a ROM within the computational unit 102 (a plurality of types of dictionary data is stored). Because dictionary data exists for every type of object and specific area, objects of different types can be detected, by switching dictionary data… a detector that performs object detection by CNN (convolutional neural network) is used for the detector 213, and dictionary data that is used by the detector in detecting each object is a learned parameter generated in advance using machine learning of a CNN… For example, learning of a CNN by an object detection unit may be performed, by a predetermined computer performing supervised learning with image data for use in learning as an input and with the position of an object corresponding to the image data for use in learning as supervised data. Also, learning of a CNN by a dictionary estimation unit may be performed, by a predetermined computer performing supervised learning with image data for use in learning as an input and with dictionary data corresponding to an object in the image data for use in learning as supervised data. Learned parameters of a CNN are generated in the manner described above… detection of a plurality of types of objects that is fewer than the total number is applied by switching the dictionary in accordance with settings determined in advance in order to save on processing speed, bus band… the method for controlling dictionary data switching is determined, based on the calculated priorities of the respective dictionary data… the image data and dictionary data generated in step S301 are input to the detector 213, and a specific area of the object corresponding to dictionary data is detected. The dictionary data that is input depends on the dictionary data switching control method determined… In step S403, it is determined whether local dictionary data is defined for the main object. Depending on the object, an entire area or a local area that is part of the entire area is defined as the detection area… the method of setting dictionary data to be used by the detector 213 in object detection with respect to a plurality of frames… The detector 213 refers to dictionary data in order of the tables stored in the RAM, and performs detection of objects corresponding to the respective dictionary data. A table indicating the input order of dictionary data extracted to the RAM is rewritten as required according to the detection state of objects; generate dictionary data for the object detection on the basis of the training data and the information related to the network structure (e.g. apparatus for respectively detecting a plurality of different objects from an image (i.e. imaging system that performs object detection) includes learning (i.e. training) of a neural network, including a convolutional neural network (CNN), for example, and implementing a detector that performs object detection by using the convolutional neural network (i.e. a network structure of the neural network in the object detection), including setting (i.e. generating, creating, computing, calculating, etc.) dictionary data to be used by the detector in the object detection of each object (i.e. generate dictionary data for the object detection) with respect to a plurality of frames, for example, including parameters (i.e. information, data, values, etc., related to the network structure) learned using machine learning (i.e. training data) of the CNN (i.e. generate dictionary data for the object detection on the basis of the training data and the information related to the network structure), as indicated above), for example); and
dictionary data transmission unit configured to transmit the dictionary data generated by the dictionary generation unit to the imaging device (Par. [0006-7]: apparatus comprising: a storage device configured to store a plurality of dictionary data for respectively detecting a plurality of different objects from an image; and at least one processor configured to function as: a detection unit configured to use partial dictionary data of the plurality of dictionary data to detect an object corresponding to the partial dictionary data, with respect to each frame of an image of a plurality of frames obtained by an image capturing device; and a switching unit configured to switch the dictionary data to be used by the detection unit in the plurality of frames, according to a result of the object detection… apparatus comprising: an image capturing device configured to capture an image; and an image processing apparatus including: a storage device configured to store a plurality of dictionary data for respectively detecting a plurality of different objects from an image; and at least one processor configured to function as: a detection unit configured to use partial dictionary data of the plurality of dictionary data to detect an object corresponding to the partial dictionary data, with respect to each frame of an image of a plurality of frames obtained by an image capturing device; and a switching unit configured to switch the dictionary data to be used by the detection unit in the plurality of frames, according to a result of the object detection; Par. [033-36]: a program that runs on a CPU, dedicated hardware, or a combination thereof may be used. Also, the type of detection object can be changed by switching the dictionary data that is input to the detector 213… Dictionary data is data in which features of corresponding objects are registered, for example, and control commands for the logic circuits are defined for every object type… dictionary data for every object is stored by a ROM within the computational unit 102 (a plurality of types of dictionary data is stored). Because dictionary data exists for every type of object and specific area, objects of different types can be detected, by switching dictionary data. For example, the dictionary data priority calculation unit 211 calculates the priority of dictionary data every predetermined period such as every frame, and, based on the computation result, the dictionary data switching control unit 212 determines dictionary data to be input to the detector 213… a detector that performs object detection by CNN (convolutional neural network) is used for the detector 213, and dictionary data that is used by the detector in detecting each object is a learned parameter generated in advance using machine learning of a CNN in an external device (PC) or the image capturing apparatus 100… Machine learning of a CNN can be performed by any technique. For example, a predetermined computer such as a server may perform machine learning of a CNN, and the image capturing apparatus 100 may acquire the learned CNN from the predetermined computer; and transmit the dictionary data generated to the imaging device (e.g. apparatus for respectively detecting a plurality of different objects from an image (i.e. imaging system that performs object detection) includes learning (i.e. training) of a neural network, including a convolutional neural network (CNN), for example, and implementing a detector that performs object detection by using the convolutional neural network (i.e. a network structure of the neural network in the object detection), including setting (i.e. generating, creating, computing, calculating, etc.) dictionary data to be used by the detector in the object detection of each object (i.e. the dictionary data generated) with respect to a plurality of frames obtained by an image capturing device, including a camera, for example, including a plurality of types of dictionary data stored by a ROM within a computational unit, such as a storage device configured to store a plurality of dictionary data for respectively detecting a plurality of different objects from an image and at least one processor configured to switch the dictionary data to be used by the detection (i.e. and transmit the dictionary data generated to the imaging device), as indicated above), for example),
wherein the information related to the network structure includes information related to at least one of an image size of input data, a number of channels of the input data, a number of parameters of the neural network, a memory capacity, a type of a layer and a type of an activation function, and a product-sum operation specification (Par. [0035-36]: a detector that performs object detection by CNN (convolutional neural network) is used for the detector 213, and dictionary data that is used by the detector in detecting each object is a learned parameter generated in advance using machine learning of a CNN in an external device (PC) or the image capturing apparatus 100… Learned parameters of a CNN are generated in the manner described above. Learning of a CNN may be performed by the image capturing apparatus 100 or the image processing apparatus described above; wherein the information related to the network structure includes information related to at least one of an image size of input data, a number of channels of the input data, a number of parameters of the neural network, a memory capacity, a type of a layer and a type of an activation function, and a product-sum operation specification (e.g. apparatus for respectively detecting a plurality of different objects from an image (i.e. imaging system that performs object detection) includes learning (i.e. training) of a neural network (i.e. the network structure), including a convolutional neural network (CNN), for example, and implementing a detector that performs object detection by using the convolutional neural network, including parameters (i.e. a number of parameters of the neural network) learned by using machine learning of the CNN (i.e. wherein the information related to the network structure includes information related to at least one of a number of parameters of the neural network), and performs object detection by using the convolutional neural network as indicated above), for example).
Regarding claim 20, claim 19 is incorporated and Kashu discloses the processing server (Par. [0006]), wherein the dictionary generation unit is configured to pick up a dictionary suitable for an object of the training data from a plurality of pieces of the dictionary data prepared in advance (Par. [0035-36]: a detector that performs object detection by CNN (convolutional neural network) is used for the detector 213, and dictionary data that is used by the detector in detecting each object is a learned parameter generated in advance using machine learning of a CNN in an external device (PC) or the image capturing apparatus 100… Machine learning of a CNN can be performed by any technique. For example, a predetermined computer such as a server may perform machine learning of a CNN, and the image capturing apparatus 100 may acquire the learned CNN from the predetermined computer; Par. [0064-65]: detection cycles of respective dictionary data that were determined in the previous frame are initialized. Here, the detection cycle is a parameter [frame/detection] representing the number of frames per detection result acquisition… For example, a plurality of methods may be stored in advance in the ROM within the computational unit 102, in the form of data tables indicating the combinations of dictionary data to be used by the detector 213 in respective frames. The control unit 201 selects a table according to settings (people priority, animal (dog, cat, bird) priority, vehicle (two-wheeled, four-wheeled) priority, etc.) configured by the user as to which object to detect as object detection, for example, and extracts the selected table to the RAM… If there is such dictionary data, the processing advances to step S502, and the detection cycle of the dictionary thereof is set to 1 [frame/detection] (select as dictionary data for acquiring per frame detection results).
Regarding claim 21, claim 19 is incorporated and Kashu discloses the processing server (Par. [0006]), wherein the dictionary generation unit is configured to generate the dictionary data by performing learning on the basis of the training data (Par. [0003]: detection of specific objects is performed by inputting images to a detector together with dictionary data obtained by learning the object to be detected. By changing the dictionary data input to the detector, it is possible to detect different types of objects from within an image; Par. [0033-64]: detector 213 performs processing for detecting a specific area (e.g., person's face or pupil, dog's face or pupil) from an image… As for the detection technique, any known method such as AdaBoost or a convolutional neural network are be used… Dictionary data is data in which features of corresponding objects are registered, for example, and control commands for the logic circuits are defined for every object type… dictionary data for every object is stored by a ROM within the computational unit 102 (a plurality of types of dictionary data is stored). Because dictionary data exists for every type of object and specific area, objects of different types can be detected, by switching dictionary data… a detector that performs object detection by CNN (convolutional neural network) is used for the detector 213, and dictionary data that is used by the detector in detecting each object is a learned parameter generated in advance using machine learning of a CNN… For example, learning of a CNN by an object detection unit may be performed, by a predetermined computer performing supervised learning with image data for use in learning as an input and with the position of an object corresponding to the image data for use in learning as supervised data. Also, learning of a CNN by a dictionary estimation unit may be performed, by a predetermined computer performing supervised learning with image data for use in learning as an input and with dictionary data corresponding to an object in the image data for use in learning as supervised data. Learned parameters of a CNN are generated in the manner described above… detection of a plurality of types of objects that is fewer than the total number is applied by switching the dictionary in accordance with settings determined in advance in order to save on processing speed, bus band… the method for controlling dictionary data switching is determined, based on the calculated priorities of the respective dictionary data… the image data and dictionary data generated in step S301 are input to the detector 213, and a specific area of the object corresponding to dictionary data is detected. The dictionary data that is input depends on the dictionary data switching control method determined… In step S403, it is determined whether local dictionary data is defined for the main object. Depending on the object, an entire area or a local area that is part of the entire area is defined as the detection area… the method of setting dictionary data to be used by the detector 213 in object detection with respect to a plurality of frames… The detector 213 refers to dictionary data in order of the tables stored in the RAM, and performs detection of objects corresponding to the respective dictionary data. A table indicating the input order of dictionary data extracted to the RAM is rewritten as required according to the detection state of objects; generate the dictionary data by performing learning on the basis of the training data (e.g. apparatus for respectively detecting a plurality of different objects from an image includes learning (i.e. training) of a neural network, including a convolutional neural network, and implementing a detector that performs object detection by using the convolutional neural network (i.e. a network structure of the neural network in the object detection), for example, including setting dictionary data to be used by the detector in object detection (i.e. generate dictionary data for the object detection) with respect to a plurality of frames, for example, including dictionary data that is used by the object detector in detecting each object as a learned parameter generated in advance using machine learning of the CNN (i.e. generate the dictionary data by performing learning on the basis of the training data), as indicated above), for example).
Regarding claim 22, claim 19 is incorporated and Kashu discloses the processing server (Par. [0006]), wherein the training data and the information related to the network structure are acquired from the imaging device or an information processing terminal that is different from the imaging device (Par. [0035-36]: detector that performs object detection by CNN (convolutional neural network) is used for the detector 213, and dictionary data that is used by the detector in detecting each object is a learned parameter generated in advance using machine learning of a CNN… Machine learning of a CNN can be performed by any technique. For example, a predetermined computer such as a server may perform machine learning of a CNN, and the image capturing apparatus 100 may acquire the learned CNN from the predetermined computer… For example, learning of a CNN by an object detection unit may be performed, by a predetermined computer performing supervised learning with image data for use in learning as an input… Also, learning of a CNN by a dictionary estimation unit may be performed, by a predetermined computer performing supervised learning with image data for use in learning as an input and with dictionary data corresponding to an object in the image data for use in learning as supervised data. Learned parameters of a CNN are generated in the manner described above. Learning of a CNN may be performed by the image capturing apparatus 100 or the image processing apparatus described above).
Regarding claim 24, is a corresponding method claim rejected as applied to the apparatus claim 1 above.
Regarding claim 25, is a corresponding method claim rejected as applied to the imaging device claim 16 above.
Regarding claim 26, is a corresponding method claim rejected as applied to the processing server claim 19 above.
Regarding claim 27, is a corresponding computer readable medium claim rejected as applied to the apparatus claim 1 above.
Regarding claim 28, is a corresponding computer readable medium claim rejected as applied to the apparatus claim 16 above.
Regarding claim 29, is a corresponding computer readable medium claim rejected as applied to the apparatus claim 19 above.
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claims 10 and 30-31 are rejected under 35 U.S.C. 103 as being unpatentable over Kashu, as applied to claims 1, 16, and 19 above, in view of SUDO et al. (US PG Publication No. 2019/0260924 A1), hereafter referred to as SUDO.
Regarding claim 10, claim 1 is incorporated and Kashu discloses the imaging system (Par. [0006]), but fails to teach the following as further recited in claim 10.
However, SUDO teaches wherein the network structure designation unit is configured to designate the network structure by designating a model of the imaging device (Par. [0020] FIG. 7 is a diagram showing a display example of an inference model selection screen (image dictionary selection screen) of the display unit of the image pickup apparatus; Par. [0035-48]: control unit 20 has multiple control sections (21 to 24) controlling various types of constituents that configure the image pickup apparatus 1, an image processing section 25 that processes the image signal, and other components. The control unit 20 has centralized control over the multiple control sections (21 to 24), the image processing section 25, and other components, and thereby functions as an image pickup control unit that controls the image pickup operation of the image pickup apparatus 1. Here, the control unit 20 as the image pickup control unit performs image pickup control on the basis of an image signal acquired by the image pickup unit 10 and multiple target image dictionaries (inference models) stored in a storage section … inference engine 30 is configured of an electronic circuit or program software that makes a predetermined inference (to be described later in detail) on a main object (image pickup target) included in an image displayed by an image signal acquired by the image pickup unit 10, on the basis of the image signal and inference model data (also referred to as target image dictionary) generated beforehand by the external equipment 100, for example. The inference engine 30 also performs processing such as determining a specific target type on the basis of an image signal acquired by the image pickup unit 10 and multiple target image dictionaries stored in the storage section 31, and selecting a target image dictionary corresponding to the determined specific target type from among the multiple target image dictionaries… the storage section 31 is a constituent section that stores multiple target image dictionaries (inference models)… the inference model is generated by extracting a feature part of a predetermined target (object) by using machine learning… Multiple inference models generated in the external equipment 100 are stored in a storage unit (not shown) inside the external equipment 100. The image pickup apparatus 1 is configured to perform data communication with the external equipment 100 through the communication unit 41 as needed, to read out a desired inference model from the storage unit (not shown) of the external equipment 100, and store the inference model in the storage section 31 of the inference engine 30 to use the inference model when necessary; Par. [0071-80]: learning unit 101 is configured of an electronic circuit or program software having a function of creating an inference model (target image dictionary). The learning unit 101 is configured to include a population creation section 102, an output setting section 103, an input-output modeling section 104, a communication section 105, and other components… The input-output modeling section 104 is an electronic circuit or program software that performs modeling processing on the basis of multiple pieces of image data included in the image population created by the population creation section 102 and various types of information set by the output setting section 103, and outputs the processing result as a learning model; Par. [0100-108]: a brief description of the effect of the image pickup apparatus 1 will be given… When the user desires to pick up an image of a specific image pickup target (e.g., “bird”) by using the image pickup apparatus 1 previously provided with inference models, first, an inference model (bird dictionary, see FIG. 7) corresponding to “bird” which is the desired image pickup target is set to be used in the image pickup apparatus… in order to pick up an image of a desired image pickup target such as “bird”, a corresponding inference model (bird dictionary) is set to be used in the image pickup apparatus 1. When the desired image pickup target (“bird” in this case) is captured in a live view image, the image pickup target (“bird” in this case) is detected on the basis of the image data of the live view image and the inference model (bird dictionary). Additionally, an image pickup parameter appropriate for the image pickup target “bird” is automatically set according to the surrounding environment. Then, the image pickup operation is automatically performed; Par. [0122-123]: an inference model having high reliability on the main object of the live view image is automatically selected, and the selected inference model is set to be used… Information on the inference model thus set automatically is superimposed on a live view image displayed on the display screen of the display unit 43 as shown in FIG. 6, for example; designate the network structure by designating a model of the imaging device (e.g. image pickup apparatus (i.e. imaging device) includes inference models (i.e. models of the imaging device) that are generated (i.e. designated, determined, etc.) by extracting a feature part of a predetermined target object and by using machine learning(i.e. network structure), for example, on the basis of an image signal and generated inference model data, also referred to as target image dictionary (i.e. designate the network structure by designating a model of the imaging device), as indicated above), for example).
Kashu and SUDO are considered to be analogous art because they pertain to image processing applications. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to modify the apparatus for respectively detecting a plurality of different objects from an image includes learning of a neural network, including a convolutional neural network, and implementing a detector that performs object detection by using the convolutional neural network (as disclosed by Kashu) with designate the network structure by designating a model of the imaging device (as taught by SUDO, Abstract, Par. [0020, 35-49]) to recognize a target in a picked up image, or an image obtained in other ways, by using machine learning (i.e. the network structure), for example, by using an inference model used for making an inference on a target (object) included in an image data, the image data including an unknown target (object), for example, by selecting a target image dictionary corresponding to a determined specific target type from among the multiple target image dictionaries (SUDO, Abstract, Par. [0020, 28, 35-49, 157]).
Regarding claim 30, claim 1 is incorporated and Kashu discloses the imaging system (Par. [0006]), but fails to teach the following as further recited in claim 30.
However, SUDO teaches wherein the network structure designation unit is configured to designate the network structure by designating a model of the imaging device (Par. [0020] FIG. 7 is a diagram showing a display example of an inference model selection screen (image dictionary selection screen) of the display unit of the image pickup apparatus; Par. [0035-48]: control unit 20 has multiple control sections (21 to 24) controlling various types of constituents that configure the image pickup apparatus 1, an image processing section 25 that processes the image signal, and other components. The control unit 20 has centralized control over the multiple control sections (21 to 24), the image processing section 25, and other components, and thereby functions as an image pickup control unit that controls the image pickup operation of the image pickup apparatus 1. Here, the control unit 20 as the image pickup control unit performs image pickup control on the basis of an image signal acquired by the image pickup unit 10 and multiple target image dictionaries (inference models) stored in a storage section … inference engine 30 is configured of an electronic circuit or program software that makes a predetermined inference (to be described later in detail) on a main object (image pickup target) included in an image displayed by an image signal acquired by the image pickup unit 10, on the basis of the image signal and inference model data (also referred to as target image dictionary) generated beforehand by the external equipment 100, for example. The inference engine 30 also performs processing such as determining a specific target type on the basis of an image signal acquired by the image pickup unit 10 and multiple target image dictionaries stored in the storage section 31, and selecting a target image dictionary corresponding to the determined specific target type from among the multiple target image dictionaries… the storage section 31 is a constituent section that stores multiple target image dictionaries (inference models)… the inference model is generated by extracting a feature part of a predetermined target (object) by using machine learning… Multiple inference models generated in the external equipment 100 are stored in a storage unit (not shown) inside the external equipment 100. The image pickup apparatus 1 is configured to perform data communication with the external equipment 100 through the communication unit 41 as needed, to read out a desired inference model from the storage unit (not shown) of the external equipment 100, and store the inference model in the storage section 31 of the inference engine 30 to use the inference model when necessary; Par. [0071-80]: learning unit 101 is configured of an electronic circuit or program software having a function of creating an inference model (target image dictionary). The learning unit 101 is configured to include a population creation section 102, an output setting section 103, an input-output modeling section 104, a communication section 105, and other components… The input-output modeling section 104 is an electronic circuit or program software that performs modeling processing on the basis of multiple pieces of image data included in the image population created by the population creation section 102 and various types of information set by the output setting section 103, and outputs the processing result as a learning model; Par. [0100-108]: a brief description of the effect of the image pickup apparatus 1 will be given… When the user desires to pick up an image of a specific image pickup target (e.g., “bird”) by using the image pickup apparatus 1 previously provided with inference models, first, an inference model (bird dictionary, see FIG. 7) corresponding to “bird” which is the desired image pickup target is set to be used in the image pickup apparatus… in order to pick up an image of a desired image pickup target such as “bird”, a corresponding inference model (bird dictionary) is set to be used in the image pickup apparatus 1. When the desired image pickup target (“bird” in this case) is captured in a live view image, the image pickup target (“bird” in this case) is detected on the basis of the image data of the live view image and the inference model (bird dictionary). Additionally, an image pickup parameter appropriate for the image pickup target “bird” is automatically set according to the surrounding environment. Then, the image pickup operation is automatically performed; Par. [0122-123]: an inference model having high reliability on the main object of the live view image is automatically selected, and the selected inference model is set to be used… Information on the inference model thus set automatically is superimposed on a live view image displayed on the display screen of the display unit 43 as shown in FIG. 6, for example; designate the network structure by designating a model of the imaging device (e.g. image pickup apparatus (i.e. imaging device) includes inference models (i.e. models of the imaging device) that are generated (i.e. designated, determined, etc.) by extracting a feature part of a predetermined target object and by using machine learning(i.e. network structure), for example, on the basis of an image signal and generated inference model data, also referred to as target image dictionary (i.e. designate the network structure by designating a model of the imaging device), as indicated above), for example), wherein the information related to a [the] network structure differs depending on a model name or model ID of the imaging device (Par. [0100-108]: a brief description of the effect of the image pickup apparatus 1 will be given… When the user desires to pick up an image of a specific image pickup target (e.g., “bird”) by using the image pickup apparatus 1 previously provided with inference models, first, an inference model (bird dictionary, see FIG. 7) corresponding to “bird” which is the desired image pickup target is set to be used in the image pickup apparatus… in order to pick up an image of a desired image pickup target such as “bird”, a corresponding inference model (bird dictionary) is set to be used in the image pickup apparatus 1. When the desired image pickup target (“bird” in this case) is captured in a live view image, the image pickup target (“bird” in this case) is detected on the basis of the image data of the live view image and the inference model (bird dictionary). Additionally, an image pickup parameter appropriate for the image pickup target “bird” is automatically set according to the surrounding environment. Then, the image pickup operation is automatically performed; Par. [0127-128]: an inference model selection screen (image dictionary selection screen) as shown in FIG. 7 is displayed on the display screen of the display unit 43. Here, FIG. 7 is an example of screen display when the operation mode of the image pickup apparatus is set to setting mode, and exemplifies a state where an inference model selection and setting screen is displayed… In the inference model selection and setting screen, as shown in FIG. 7, multiple inference models (target image dictionaries) previously stored in the storage section 31 of the inference engine 30 of the image pickup apparatus 1 are displayed in a list (reference numeral 203 in FIG. 7). The example shown in FIG. 7 exemplifies a state where names, icons, or the like are displayed to indicate each of the inference models… For example… the user touches an icon or the like displaying a desired inference model from the list displayed on the display screen with his/her finger… Then, the displayed icon of the selected and set inference model changes to a display form visibly recognizable by flashing or emphasizing, for example. In the example shown in FIG. 7, as a display example of emphasizing the selected inference model, a frame surrounding the characters of the name indicating the inference model is displayed in a bold frame; Par. [0137-142]: assume that the user has captured a desired image pickup target 201 (object) within image pickup range F by use of the image pickup apparatus 1, for example. The live view image displayed on the display unit 43 at this time is as shown in FIG. 6, for example. At this time, the “bird” inference model is set in the image pickup apparatus 1… In this state, the image pickup apparatus 1 performs control processing of step S18 on the basis of the inference result of the inference processing in the aforementioned step S15. The control processing includes control such as a focusing operation of identifying “bird” which is the main object in the live view image, and focusing on the “bird”… if the user recaptures “bird” which is the image pickup target 201 after moving as in FIG. 10 by holding up the image pickup apparatus 1 again, for example, the inference using the maintained inference model is continued; wherein the information related to the network structure differs depending on a model name or model ID of the imaging device (e.g. image pickup apparatus (i.e. imaging device) includes inference models (i.e. models of the imaging device) that are generated (i.e. designated, determined, etc.) by extracting a feature part of a predetermined target object and by using machine learning(i.e. network structure), for example, on the basis of an image signal and generated inference model data (i.e. the information related to the network structure), also referred to as target image dictionary, for example, including names or icons that are displayed to indicate each of the inference models on the image pickup apparatus (i.e. wherein the information related to the network structure differs depending on a model name or model ID of the imaging device), as indicated above), for example).
The same motivation to combine above-mentioned teachings applies, as previously indicated in claim 10.
Regarding claim 31, claim 16 is incorporated and Kashu discloses the imaging device (Par. [0006]), but fails to teach the following as further recited in claim 31.
However, SUDO teaches wherein the network structure designation unit is configured to designate the network structure by designating a model of the imaging device (Par. [0020] FIG. 7 is a diagram showing a display example of an inference model selection screen (image dictionary selection screen) of the display unit of the image pickup apparatus; Par. [0035-48]: control unit 20 has multiple control sections (21 to 24) controlling various types of constituents that configure the image pickup apparatus 1, an image processing section 25 that processes the image signal, and other components. The control unit 20 has centralized control over the multiple control sections (21 to 24), the image processing section 25, and other components, and thereby functions as an image pickup control unit that controls the image pickup operation of the image pickup apparatus 1. Here, the control unit 20 as the image pickup control unit performs image pickup control on the basis of an image signal acquired by the image pickup unit 10 and multiple target image dictionaries (inference models) stored in a storage section … inference engine 30 is configured of an electronic circuit or program software that makes a predetermined inference (to be described later in detail) on a main object (image pickup target) included in an image displayed by an image signal acquired by the image pickup unit 10, on the basis of the image signal and inference model data (also referred to as target image dictionary) generated beforehand by the external equipment 100, for example. The inference engine 30 also performs processing such as determining a specific target type on the basis of an image signal acquired by the image pickup unit 10 and multiple target image dictionaries stored in the storage section 31, and selecting a target image dictionary corresponding to the determined specific target type from among the multiple target image dictionaries… the storage section 31 is a constituent section that stores multiple target image dictionaries (inference models)… the inference model is generated by extracting a feature part of a predetermined target (object) by using machine learning… Multiple inference models generated in the external equipment 100 are stored in a storage unit (not shown) inside the external equipment 100. The image pickup apparatus 1 is configured to perform data communication with the external equipment 100 through the communication unit 41 as needed, to read out a desired inference model from the storage unit (not shown) of the external equipment 100, and store the inference model in the storage section 31 of the inference engine 30 to use the inference model when necessary; Par. [0071-80]: learning unit 101 is configured of an electronic circuit or program software having a function of creating an inference model (target image dictionary). The learning unit 101 is configured to include a population creation section 102, an output setting section 103, an input-output modeling section 104, a communication section 105, and other components… The input-output modeling section 104 is an electronic circuit or program software that performs modeling processing on the basis of multiple pieces of image data included in the image population created by the population creation section 102 and various types of information set by the output setting section 103, and outputs the processing result as a learning model; Par. [0100-108]: a brief description of the effect of the image pickup apparatus 1 will be given… When the user desires to pick up an image of a specific image pickup target (e.g., “bird”) by using the image pickup apparatus 1 previously provided with inference models, first, an inference model (bird dictionary, see FIG. 7) corresponding to “bird” which is the desired image pickup target is set to be used in the image pickup apparatus… in order to pick up an image of a desired image pickup target such as “bird”, a corresponding inference model (bird dictionary) is set to be used in the image pickup apparatus 1. When the desired image pickup target (“bird” in this case) is captured in a live view image, the image pickup target (“bird” in this case) is detected on the basis of the image data of the live view image and the inference model (bird dictionary). Additionally, an image pickup parameter appropriate for the image pickup target “bird” is automatically set according to the surrounding environment. Then, the image pickup operation is automatically performed; Par. [0122-123]: an inference model having high reliability on the main object of the live view image is automatically selected, and the selected inference model is set to be used… Information on the inference model thus set automatically is superimposed on a live view image displayed on the display screen of the display unit 43 as shown in FIG. 6, for example; designate the network structure by designating a model of the imaging device (e.g. image pickup apparatus (i.e. imaging device) includes inference models (i.e. models of the imaging device) that are generated (i.e. designated, determined, etc.) by extracting a feature part of a predetermined target object and by using machine learning(i.e. network structure), for example, on the basis of an image signal and generated inference model data, also referred to as target image dictionary (i.e. designate the network structure by designating a model of the imaging device), as indicated above), for example), wherein the information related to a [the] network structure differs depending on a model name or model ID of the imaging device (Par. [0100-108]: a brief description of the effect of the image pickup apparatus 1 will be given… When the user desires to pick up an image of a specific image pickup target (e.g., “bird”) by using the image pickup apparatus 1 previously provided with inference models, first, an inference model (bird dictionary, see FIG. 7) corresponding to “bird” which is the desired image pickup target is set to be used in the image pickup apparatus… in order to pick up an image of a desired image pickup target such as “bird”, a corresponding inference model (bird dictionary) is set to be used in the image pickup apparatus 1. When the desired image pickup target (“bird” in this case) is captured in a live view image, the image pickup target (“bird” in this case) is detected on the basis of the image data of the live view image and the inference model (bird dictionary). Additionally, an image pickup parameter appropriate for the image pickup target “bird” is automatically set according to the surrounding environment. Then, the image pickup operation is automatically performed; Par. [0127-128]: an inference model selection screen (image dictionary selection screen) as shown in FIG. 7 is displayed on the display screen of the display unit 43. Here, FIG. 7 is an example of screen display when the operation mode of the image pickup apparatus is set to setting mode, and exemplifies a state where an inference model selection and setting screen is displayed… In the inference model selection and setting screen, as shown in FIG. 7, multiple inference models (target image dictionaries) previously stored in the storage section 31 of the inference engine 30 of the image pickup apparatus 1 are displayed in a list (reference numeral 203 in FIG. 7). The example shown in FIG. 7 exemplifies a state where names, icons, or the like are displayed to indicate each of the inference models… For example… the user touches an icon or the like displaying a desired inference model from the list displayed on the display screen with his/her finger… Then, the displayed icon of the selected and set inference model changes to a display form visibly recognizable by flashing or emphasizing, for example. In the example shown in FIG. 7, as a display example of emphasizing the selected inference model, a frame surrounding the characters of the name indicating the inference model is displayed in a bold frame; Par. [0137-142]: assume that the user has captured a desired image pickup target 201 (object) within image pickup range F by use of the image pickup apparatus 1, for example. The live view image displayed on the display unit 43 at this time is as shown in FIG. 6, for example. At this time, the “bird” inference model is set in the image pickup apparatus 1… In this state, the image pickup apparatus 1 performs control processing of step S18 on the basis of the inference result of the inference processing in the aforementioned step S15. The control processing includes control such as a focusing operation of identifying “bird” which is the main object in the live view image, and focusing on the “bird”… if the user recaptures “bird” which is the image pickup target 201 after moving as in FIG. 10 by holding up the image pickup apparatus 1 again, for example, the inference using the maintained inference model is continued; wherein the information related to the network structure differs depending on a model name or model ID of the imaging device (e.g. image pickup apparatus (i.e. imaging device) includes inference models (i.e. models of the imaging device) that are generated (i.e. designated, determined, etc.) by extracting a feature part of a predetermined target object and by using machine learning(i.e. network structure), for example, on the basis of an image signal and generated inference model data (i.e. the information related to the network structure), also referred to as target image dictionary, for example, including names or icons that are displayed to indicate each of the inference models on the image pickup apparatus (i.e. wherein the information related to the network structure differs depending on a model name or model ID of the imaging device), as indicated above), for example).
The same motivation to combine above-mentioned teachings applies, as previously indicated in claim 10.
Regarding claim 32, claim 19 is incorporated and Kashu discloses the information processing server (Par. [0006, 35-36]), but fails to teach the following as further recited in claim 30.
However, SUDO teaches wherein the information related to a [the] network structure differs depending on a model name or model ID of the imaging device (Par. [0100-108]: a brief description of the effect of the image pickup apparatus 1 will be given… When the user desires to pick up an image of a specific image pickup target (e.g., “bird”) by using the image pickup apparatus 1 previously provided with inference models, first, an inference model (bird dictionary, see FIG. 7) corresponding to “bird” which is the desired image pickup target is set to be used in the image pickup apparatus… in order to pick up an image of a desired image pickup target such as “bird”, a corresponding inference model (bird dictionary) is set to be used in the image pickup apparatus 1. When the desired image pickup target (“bird” in this case) is captured in a live view image, the image pickup target (“bird” in this case) is detected on the basis of the image data of the live view image and the inference model (bird dictionary). Additionally, an image pickup parameter appropriate for the image pickup target “bird” is automatically set according to the surrounding environment. Then, the image pickup operation is automatically performed; Par. [0127-128]: an inference model selection screen (image dictionary selection screen) as shown in FIG. 7 is displayed on the display screen of the display unit 43. Here, FIG. 7 is an example of screen display when the operation mode of the image pickup apparatus is set to setting mode, and exemplifies a state where an inference model selection and setting screen is displayed… In the inference model selection and setting screen, as shown in FIG. 7, multiple inference models (target image dictionaries) previously stored in the storage section 31 of the inference engine 30 of the image pickup apparatus 1 are displayed in a list (reference numeral 203 in FIG. 7). The example shown in FIG. 7 exemplifies a state where names, icons, or the like are displayed to indicate each of the inference models… For example… the user touches an icon or the like displaying a desired inference model from the list displayed on the display screen with his/her finger… Then, the displayed icon of the selected and set inference model changes to a display form visibly recognizable by flashing or emphasizing, for example. In the example shown in FIG. 7, as a display example of emphasizing the selected inference model, a frame surrounding the characters of the name indicating the inference model is displayed in a bold frame; Par. [0137-142]: assume that the user has captured a desired image pickup target 201 (object) within image pickup range F by use of the image pickup apparatus 1, for example. The live view image displayed on the display unit 43 at this time is as shown in FIG. 6, for example. At this time, the “bird” inference model is set in the image pickup apparatus 1… In this state, the image pickup apparatus 1 performs control processing of step S18 on the basis of the inference result of the inference processing in the aforementioned step S15. The control processing includes control such as a focusing operation of identifying “bird” which is the main object in the live view image, and focusing on the “bird”… if the user recaptures “bird” which is the image pickup target 201 after moving as in FIG. 10 by holding up the image pickup apparatus 1 again, for example, the inference using the maintained inference model is continued; wherein the information related to the network structure differs depending on a model name or model ID of the imaging device (e.g. image pickup apparatus (i.e. imaging device) includes inference models (i.e. models of the imaging device) that are generated (i.e. designated, determined, etc.) by extracting a feature part of a predetermined target object and by using machine learning(i.e. network structure), for example, on the basis of an image signal and generated inference model data (i.e. the information related to the network structure), also referred to as target image dictionary, for example, including names or icons that are displayed to indicate each of the inference models on the image pickup apparatus (i.e. wherein the information related to the network structure differs depending on a model name or model ID of the imaging device), as indicated above), for example).
The same motivation to combine above-mentioned teachings applies, as previously indicated in claim 10.
Claims 11 and 13-14 are rejected under 35 U.S.C. 103 as being unpatentable over Kashu, as applied to claim 1 above, in view of Fukuda et al. (US PG Publication No. 2011/0091072 A1), hereafter referred to as Fukuda, Applicant cited prior art furnished via IDS as JP 2011-090410.
Regarding claim 11, claim 1 is incorporated and Kashu discloses the imaging system (Par. [0006]), but fails to teach the following as further recited in claim 11.
However, Fukuda teaches wherein the at least one processor or circuit is further configured to function as (Par. [0047-50]: recognition processor 109 is a component for performing an image data recognition process and detects an object of recognition from image data input by the data input unit 101… Part or all of the processing of the recognition processor 109 in FIG. 1 may be performed by software instead. In this case, the processing of the recognition processor 109 is performed by the CPU 104 executing a program stored in the ROM 105, the RAM 106, or the data storage unit 102. It is also possible to provide a general signal processor or a general image processor (not shown) so that part of this software processing of the recognition processor is performed by the general signal processor or the general image processor instead; Par. [0155]: Aspects of the present invention can also be realized by a computer of a system or apparatus (or devices such as a CPU or MPU) that reads out and executes a program recorded on a memory device to perform the functions of the above-described embodiment(s), and by a method, the steps of which are performed by a computer of a system or apparatus by, for example, reading out and executing a program recorded on a memory device to perform the functions of the above-described embodiment(s). For this purpose, the program is provided to the computer for example via a network or from a recording medium of various types serving as the memory device (for example, computer-readable medium)),
dictionary validation unit configured to validate the dictionary data generated by the dictionary generation unit (Par. [0094-130]: the data processing apparatus according to the present invention receives recognition dictionary information… the data processing apparatus 320 receives designation information about recognition dictionaries from a server… CPU 104 of the data processing apparatus 320 performs processing necessary for performing the recognition process by using the received dictionary data. To prepare for a shooting process and the like to be described later, the CPU 104 checks the number of recognition dictionaries stored in the data processing apparatus 320… The CPU 104 generates a list of the recognition dictionaries, which is also referred to in the shooting process… processing of determining whether to terminate the iterative process is performed. NumOfActiveDic is a variable representing the number of valid dictionaries held in the data processing apparatus 320… the CPU 104 identifies an i-th dictionary from the list of received recognition dictionaries. The recognition processor 109 uses the identified recognition dictionary to perform the recognition process for an object of shooting in the input image data… The CPU 104 stores the recognition result in the recognition process step in step S705 in the RAM 106 or the data storage unit 102 in a form that allows knowing which recognition dictionary was used for the detection; validate the dictionary data generated (e.g. data processing apparatus comprising recognition processor for performing an image data recognition process and detect an object of recognition from image data input includes a CPU that checks a number of recognition dictionaries stored in the data processing apparatus, including a number of valid dictionaries held in the data processing apparatus (i.e. validate the dictionary data generated), as indicated above), for example),
wherein in a case where the dictionary data has been validated by the dictionary validation unit, the imaging device is configured to perform the predetermined imaging control on the object detected through the object detection (Par. [0100-130]: processing of determining whether to terminate the iterative process is performed. NumOfActiveDic is a variable representing the number of valid dictionaries held in the data processing apparatus 320… above iterative process is repeated for the number of valid dictionaries held in the data processing apparatus 320… recognition process takes, as an input, the image data input through the data input unit 101 and stored in the RAM 106 in step S702. According to the process described in the flow diagram of FIG. 6, the CPU 104 identifies an i-th dictionary from the list of received recognition dictionaries. The recognition processor 109 uses the identified recognition dictionary to perform the recognition process for an object of shooting in the input image data… The CPU 104 stores the recognition result in the recognition process step in step S705 in the RAM 106 or the data storage unit 102 in a form that allows knowing which recognition dictionary was used for the detection… at a location where the data processing apparatus inputs data, necessary recognition dictionaries are distributed from the server to the data processing apparatus, and the distributed recognition dictionaries are used to perform recognition upon input of image data… the use of recognition dictionaries can be controlled according to a condition desired by the server that provides the recognition dictionaries; wherein in a case where the dictionary data has been validated, the imaging device is configured to perform the predetermined imaging control on the object detected through the object detection (e.g. data processing apparatus comprising recognition processor for performing an image data recognition process and detect an object of recognition from image data input includes a CPU that checks a number of recognition dictionaries stored in the data processing apparatus, including a number of valid dictionaries held in the data processing apparatus (i.e. wherein in a case where the dictionary data has been validated), for example, and the recognition dictionaries are used to perform recognition upon input of image data (i.e. the imaging device is configured to perform the predetermined imaging control on the object detected through the object detection), as indicated above), for example), and
in a case where the dictionary data has not been validated by the dictionary validation unit, the imaging device is configured not to perform the predetermined imaging control (Par. [0115-130]: before the process flow in step S701, a process of a recognition dictionary check step in step S801 is performed. The process of the recognition dictionary check step will be described with reference to a flow diagram in FIG. 9… In step S901, for each recognition dictionary in the data processing apparatus 320, the CPU 104 performs a determination process according to an invalidation determination condition… the server has a plurality of types of recognition dictionaries and controls an invalidation condition on the recognition dictionaries. FIG. 10 is a diagram showing a configuration of a recognition dictionary used in the third embodiment… The invalidation of recognition dictionaries performed in the second embodiment is also performed in the third embodiment, but in this case the server performs the invalidation… The recognition dictionaries do not need to have a uniform invalidation condition, so that the invalidation condition may vary among the recognition dictionaries. In this case, the data processing apparatus is configured to be capable of addressing any of the several invalidation conditions listed; and in a case where the dictionary data has not been validated, the imaging device is configured not to perform the predetermined imaging control (e.g. data processing apparatus comprising recognition processor for performing an image data recognition process and detect an object of recognition from image data input includes a CPU that checks a number of recognition dictionaries stored in the data processing apparatus and performs a determination process according to an invalidation (i.e. the dictionary data has not been validated) determination condition (i.e. and in a case where the dictionary data has not been validated, the imaging device is configured not to perform the predetermined imaging control), as indicated above), for example).
Kashu and Fukuda are considered to be analogous art because they pertain to image processing applications. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to modify the apparatus for respectively detecting a plurality of different objects from an image includes learning of a neural network, including a convolutional neural network, and implementing a detector that performs object detection by using the convolutional neural network (as disclosed by Kashu) with wherein the at least one processor or circuit is further configured to function as, dictionary validation unit configured to validate the dictionary data generated by the dictionary generation unit, wherein in a case where the dictionary data has been validated by the dictionary validation unit, the imaging device is configured to perform the predetermined imaging control on the object detected through the object detection, and in a case where the dictionary data has not been validated by the dictionary validation unit, the imaging device is configured not to perform the predetermined imaging control (as taught by Fukuda, Abstract, Par. [0047-50, 94-130, 155]) to detect a specific object of shooting from digital image data, to automatically detect a particular pattern of an object of shooting from an image, to reduce the data processing load of detecting a specific object of shooting from digital image data, and to identify recognition dictionaries from among the plurality of recognition dictionaries stored in the storage unit and uses the identified recognition dictionaries to recognize the object of recognition included in the image data (Fukuda, Abstract, Par. [0002-23, 47-50, 94-130, 155]).
Regarding claim 13, claim 11 is incorporated and the combination of Kashu and Fukuda, as a whole teaches the imaging system (Kashu, Par. [0006]), wherein the dictionary validation unit is configured to validate the dictionary data through charging (Fukuda, Par. [0094-130]: the data processing apparatus according to the present invention receives recognition dictionary information… the data processing apparatus 320 receives designation information about recognition dictionaries from a server… CPU 104 of the data processing apparatus 320 performs processing necessary for performing the recognition process by using the received dictionary data. To prepare for a shooting process and the like to be described later, the CPU 104 checks the number of recognition dictionaries stored in the data processing apparatus 320… The CPU 104 generates a list of the recognition dictionaries, which is also referred to in the shooting process… processing of determining whether to terminate the iterative process is performed. NumOfActiveDic is a variable representing the number of valid dictionaries held in the data processing apparatus 320… the CPU 104 identifies an i-th dictionary from the list of received recognition dictionaries. The recognition processor 109 uses the identified recognition dictionary to perform the recognition process for an object of shooting in the input image data… The CPU 104 stores the recognition result in the recognition process step in step S705 in the RAM 106 or the data storage unit 102 in a form that allows knowing which recognition dictionary was used for the detection; Par. [0147-150]: different conditions on distributing recognition dictionaries are set for connection via the communication path 315 and connection via the communication path 1301… if recognition dictionaries are distributed to the data processing apparatus 320 via the communication path 315 provided inside or near the target facility, the user can use the recognition dictionaries free of charge. However, if recognition dictionaries are distributed to the data processing apparatus 320 via the communication path 1301, the use of the recognition dictionaries is charged and the user is billed… in a system in which recognition dictionaries can be distributed from the same server via a plurality of paths, different billing information is set for different paths. This allows increasing added values of visiting a place where the recognition dictionaries are required; validate the dictionary data through charging (e.g. data processing apparatus comprising recognition processor for performing an image data recognition process and detect an object of recognition from image data input includes a number of recognition dictionaries stored in the data processing apparatus, including a number of valid dictionaries held in the data processing apparatus, for example, and if recognition dictionaries are distributed to the data processing apparatus, the use of the recognition dictionaries is charged (i.e. validate the dictionary data through charging), as indicated above), for example).
The same motivation to combine above-mentioned teachings applies, as previously indicated in claim 11.
Regarding claim 14, claim 11 is incorporated and the combination of Kashu and Fukuda, as a whole teaches the imaging system (Kashu, Par. [0006]), wherein the dictionary validation unit is configured to validate (become available) each piece of dictionary data through charging in a case where there is a plurality of pieces of the dictionary data generated by the dictionary generation unit (Fukuda, Par. [0094-130]: the data processing apparatus according to the present invention receives recognition dictionary information… the data processing apparatus 320 receives designation information about recognition dictionaries from a server… CPU 104 of the data processing apparatus 320 performs processing necessary for performing the recognition process by using the received dictionary data. To prepare for a shooting process and the like to be described later, the CPU 104 checks the number of recognition dictionaries stored in the data processing apparatus 320… The CPU 104 generates a list of the recognition dictionaries, which is also referred to in the shooting process… processing of determining whether to terminate the iterative process is performed. NumOfActiveDic is a variable representing the number of valid dictionaries held in the data processing apparatus 320… the CPU 104 identifies an i-th dictionary from the list of received recognition dictionaries. The recognition processor 109 uses the identified recognition dictionary to perform the recognition process for an object of shooting in the input image data… The CPU 104 stores the recognition result in the recognition process step in step S705 in the RAM 106 or the data storage unit 102 in a form that allows knowing which recognition dictionary was used for the detection; Par. [0147-150]: different conditions on distributing recognition dictionaries are set for connection via the communication path 315 and connection via the communication path 1301… if recognition dictionaries are distributed to the data processing apparatus 320 via the communication path 315 provided inside or near the target facility, the user can use the recognition dictionaries free of charge. However, if recognition dictionaries are distributed to the data processing apparatus 320 via the communication path 1301, the use of the recognition dictionaries is charged and the user is billed… in a system in which recognition dictionaries can be distributed from the same server via a plurality of paths, different billing information is set for different paths. This allows increasing added values of visiting a place where the recognition dictionaries are required; validate each piece of dictionary data through charging in a case where there is a plurality of pieces of the dictionary data generated (e.g. data processing apparatus comprising recognition processor for performing an image data recognition process and detect an object of recognition from image data input includes a number of recognition dictionaries (i.e. a plurality of pieces of the dictionary data generated) stored in the data processing apparatus, including a number of valid dictionaries held in the data processing apparatus (i.e. each piece of dictionary data), for example, and if recognition dictionaries are distributed to the data processing apparatus, the use of the recognition dictionaries is charged (i.e. validate each piece of dictionary data through charging in a case where there is a plurality of pieces of the dictionary data generated), as indicated above), for example).
The same motivation to combine above-mentioned teachings applies, as previously indicated in claim 11.
Conclusion
Applicant’s amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a).
A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any extension fee pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action.
Contact
Any inquiry concerning this communication or earlier communications from the examiner should be directed to GUILLERMO RIVERA-MARTINEZ whose telephone number is 571-272-4979. The examiner can normally be reached on Monday-Friday (8am - 5pm Eastern Time). If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Andrew Bee can be reached on 571-270-5183. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of an application may be obtained from the Patent Application Information Retrieval (PAIR) system. Status information for published applications may be obtained from either Private PAIR or Public PAIR. Status information for unpublished applications is available through Private PAIR only. For more information about the PAIR system, see http://pair-direct.uspto.gov. Should you have questions on access to the Private PAIR system, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative or access to the automated information system, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/GUILLERMO M RIVERA-MARTINEZ/ Primary Examiner, Art Unit 2677