DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Priority
Receipt is acknowledged of certified copies of papers required by 37 CFR 1.55.
Information Disclosure Statement
The two information disclosure statements (IDS) filed on December 20, 204, have been considered by the examiner.
Claim Objections
Claim 1 is objected to because of the following informalities:
Claim 1 recites the limitation "the visual map" in line 4. There is insufficient antecedent basis for this limitation in the claim.
Claim 12 is objected to because of the following informalities:
Claim 11 should be amended to recite “when there is no plurality of similar scenes” for grammatical clarity.
Appropriate correction is required.
Claim Interpretation
The following is a quotation of 35 U.S.C. 112(f):
(f) Element in Claim for a Combination. – An element in a claim for a combination may be expressed as a means or step for performing a specified function without the recital of structure, material, or acts in support thereof, and such claim shall be construed to cover the corresponding structure, material, or acts described in the specification and equivalents thereof.
The following is a quotation of pre-AIA 35 U.S.C. 112, sixth paragraph:
An element in a claim for a combination may be expressed as a means or step for performing a specified function without the recital of structure, material, or acts in support thereof, and such claim shall be construed to cover the corresponding structure, material, or acts described in the specification and equivalents thereof.
The claims in this application are given their broadest reasonable interpretation using the plain meaning of the claim language in light of the specification as it would be understood by one of ordinary skill in the art. The broadest reasonable interpretation of a claim element (also commonly referred to as a claim limitation) is limited by the description in the specification when 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is invoked.
As explained in MPEP § 2181, subsection I, claim limitations that meet the following three-prong test will be interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph:
(A) the claim limitation uses the term “means” or “step” or a term used as a substitute for “means” that is a generic placeholder (also called a nonce term or a non-structural term having no specific structural meaning) for performing the claimed function;
(B) the term “means” or “step” or the generic placeholder is modified by functional language, typically, but not always linked by the transition word “for” (e.g., “means for”) or another linking word or phrase, such as “configured to” or “so that”; and
(C) the term “means” or “step” or the generic placeholder is not modified by sufficient structure, material, or acts for performing the claimed function.
Use of the word “means” (or “step”) in a claim with functional language creates a rebuttable presumption that the claim limitation is to be treated in accordance with 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph. The presumption that the claim limitation is interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is rebutted when the claim limitation recites sufficient structure, material, or acts to entirely perform the recited function.
Absence of the word “means” (or “step”) in a claim creates a rebuttable presumption that the claim limitation is not to be treated in accordance with 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph. The presumption that the claim limitation is not interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is rebutted when the claim limitation recites function without reciting sufficient structure, material or acts to entirely perform the recited function.
Claim limitations in this application that use the word “means” (or “step”) are being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, except as otherwise indicated in an Office action. Conversely, claim limitations in this application that do not use the word “means” (or “step”) are not being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, except as otherwise indicated in an Office action.
This application includes one or more claim limitations that do not use the word “means,” but are nonetheless being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, because the claim limitation(s) uses a generic placeholder that is coupled with functional language without reciting sufficient structure to perform the recited function and the generic placeholder is not preceded by a structural modifier. Such claim limitation(s) is/are:
a ‘data acquisition module’ in Claim 18,
an ‘information extraction module’ in Claim 18,
a ‘map frame determination module’ in Claim 18, and
a ‘repositioning module’ in Claim 18.
Because this/these claim limitation(s) is/are being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, it/they is/are being interpreted to cover the corresponding structure described in the specification as performing the claimed function, and equivalents thereof.
If applicant does not intend to have this/these limitation(s) interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, applicant may: (1) amend the claim limitation(s) to avoid it/them being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph (e.g., by reciting sufficient structure to perform the claimed function); or (2) present a sufficient showing that the claim limitation(s) recite(s) sufficient structure to perform the claimed function so as to avoid it/them being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph.
Each ‘module’ claimed is being interpreted as software implemented on a processor. Support for this interpretation can be found in paragraph [00132] of the instant specification.
Claim Rejections - 35 USC § 101
35 U.S.C. 101 reads as follows:
Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title.
Claim 20 is rejected under 35 U.S.C. 101 because the claimed invention is directed to non-statutory subject matter. The claim(s) does/do not fall within at least one of the four categories of patent eligible subject matter because Claims 20 are directed towards software per se.
Under the broadest reasonable interpretation, this limitation includes products that do not have a physical or tangible form (see MPEP §2106.03, “Non-limiting examples of claims that are not directed to any of the statutory categories include: Products that do not have a physical or tangible form, such as information (often referred to as “data per se”) or a computer program per se (often referred to as “software per se”) when claimed as a product without any structural recitation... For example, the BRI of machine readable media can encompass non-statutory transitory forms of signal transmission, such as a propagating electrical or electromagnetic signal per se. See In re Nuijten, 500 F.3d 1346, 84 USPQ2d 1495 (Fed. Cir. 2007). When the BRI encompasses transitory forms of signal transmission, a rejection under 35 U.S.C. 101 as failing to claim statutory subject matter would be appropriate. Thus, a claim to a computer readable medium that can be a compact disc or a carrier wave covers a non-statutory embodiment and therefore should be rejected under 35 U.S.C. 101 as being directed to non-statutory subject matter. See, e.g., Mentor Graphics v. EVE-USA, Inc., 851 F.3d at 1294-95, 112 USPQ2d at 1134 (claims to a "machine-readable medium" were non-statutory, because their scope encompassed both statutory random-access memory and non-statutory carrier waves)”).
The Examiner suggests rewriting Claim 20 to read “A non-transitory computer readable storage medium on which a computer program is stored” to ensure that Claim 20 falls under the four statutory categories of invention.
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows:
1. Determining the scope and contents of the prior art.
2. Ascertaining the differences between the prior art and the claims at issue.
3. Resolving the level of ordinary skill in the pertinent art.
4. Considering objective evidence present in the application indicating obviousness or nonobviousness.
Claims 1, 3-6, 11, 16-20 rejected under 35 U.S.C. 103 as being unpatentable over Hsieh et al. (US Pub No 2022/0301222), hereinafter Hsieh, in view of Miyatani (US Pub No 2023/0168103), hereinafter Miyatani, and further in view of Mo et al. (CN Pub No 111783838), hereinafter Mo.
As to Claim 1, Hsieh teaches a repositioning method (see Abstract, “An indoor positioning system and method are provided”), comprising:
acquiring environmental image data of a target object to be positioned at a current position (see paragraph [0047], “Step S50: configuring the image capturing device to obtain a captured image at a current position in the target area”, where the image capturing device is interpreted as the ‘target object’);
extracting features from the environmental image data (see paragraph [0048], “Step S51: configuring the computing device to input the captured image into the trained deep learning network to perform the image feature extraction on the captured image to obtain a captured image feature corresponding to the captured image”);
determining
K
map frames with a maximum similarity to the environmental image data in the visual map according to the feature information, where
K
is an integer and satisfies
K
≥
1
(see paragraph [00449], “Step S52: executing a similarity matching algorithm on the captured image feature and the plurality of virtual image features to obtain a plurality of matching virtual images with relatively high similarities to the captured image”, wherein the plurality of visual features corresponds to the map data, and a plurality is greater than 1),
and when determining that there is no symmetrical scene in the visual map for the current position of the target object according to the
K
map frames with the maximum similarity (see paragraph [0051], “After the virtual image with the highest degree of similarity is matched and obtained, in order to solve design issues related to repetitiveness and symmetry of the building, a correct image is manually selected from the plurality of matched virtual images with relatively higher similarities”, wherein symmetric image is not selected, and therefore it is determined that there is not a symmetrical scene),
and when a map frame with a maximum similarity among the K map frames is determined as a target similar scene for the current position, obtaining the position of the target object according to the map frame with the maximum similarity among the K map frames (see paragraph [0055], “After the most similar image is successfully matched, the present disclosure can utilize the most similar image to evaluate the position and the pose of the image capturing device 12 when the captured image is taken”, wherein the image capturing device 12 is the ‘target object’).
Hsieh fails to teach that there is no plurality of similar scenes in the visual map for the current position of the target object. Hsieh further fails to explicitly teach repositioning the target object according to the map frame with maximum similarity.
However, in an analogous art, Miyatani teaches a method for localizing a target object which comprises obtaining environment image data of a target object (see paragraph [0027], “Each map element is a key frame that associates an image of the environment captured by a camera mounted on the vehicle and its image feature with position information indicating the position of the camera. …The camera captures a moving image of an environment”, where the vehicle is interpreted as the target object)
and extracting features from the image data (see paragraph [0032], “feature vectors are extracted from the respective images”),
determining
K
map frames with a maximum similarity to the environmental image data in the visual map according to the feature information (see paragraph [0019], “a determination is made for each map element of interest as to whether there is another map element including an image similar to an image of the map element of interest”, wherein each ‘map element’ corresponds to a map frame, and see paragraph [0032], “More specifically, feature vectors are extracted from the respective images, and the degree of similarity of the feature vector is calculated and the degree of similarity of the image is defined by the degree similarity of the feature vector”),
and when determining that there is no plurality of similar scenes in the visual map for the current position of the target object according to the K map frames with the maximum similarity (see paragraph [0031], “In step S301, a determination is made as to similarity among the plurality of map elements acquired in step S300, thereby determining whether there is a map element including an image captured at a position where wrong position is likely to be estimated in the localization of the imaging apparatus”, where the Examiner notes that a ‘wrong position’ corresponds to similar scenes in the map elements that appears similar to a current image, but is actually located far from the current position. See paragraph [0062] of Miyatani ),
and when a map frame with a maximum similarity among the K map frames is determined as a target similar scene for the current position, repositioning the target object according to the map frame with the maximum similarity among the K map frames (see paragraph [0062], “In the localization using the SLAM technology, the localization fails if there is no map element that contains an image partially similar to an image captured at the current position. One method of recovering from this situation is to move the camera to a position corresponding to a map element where it is possible to capture an image which partially similar to another existing image, and restarting the localization at this position”).
Thus, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the determination of similar scenes taught by Miyatani with the teachings of Hsieh. The motivation for doing so would be to prevent inaccurate localization of a target object. Miyatani teaches in paragraph [0062], “However, in a case where there are a plurality of map elements including images similar to the image captured at the recovery position, there is a possibility that a map element different from a map element corresponding to the current position is selected, and thus wrong position estimation occurs in the localization at the recover position. To avoid the above situation, it is desirable to select, as the recovery start position, a position which is close to the position of the failure in the localization and which corresponds to a map element to which there is no other map element similar”).
Neither Hsieh or Miyatani explicitly teach extracting global description information from the environmental image data, or determining K map frames with a maximum similarity to the environmental image data in the visual map according to the global description information.
However, in an analogous art, Mo teaches a method for positioning a mobile robot (see paragraph [0086], “performing laser SLAM closed-loop task and repositioning task in the point cloud characteristic space by the point cloud characteristic space characterization method”) which comprises
extracting global description information from the environmental image data (see paragraph [0008], “The technical problem to be solved by the invention is to provide a point cloud feature space characterization method for laser SLAM, which can extract point cloud global description features representing scene structure information and semantic information from complex large-scene laser point clouds”),
or determining K map frames with a maximum similarity to the environmental image data in the visual map according to the global description information (see paragraph [0086], “in the actual SLAM operation, the encoder network model and parameters in step 4 are taken, global description features are extracted from the point cloud key frames, NN matching is carried out on the point cloud key frames and the global description features of the historical key frames according to Euclidean distances”).
Thus, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the network for global description information extraction taught by Mo with the teachings of Hsieh and Miyatani. The motivation for doing so would be to utilize the network in order to extract features from a large set of data (see paragraph [0040], “The invention has the beneficial effects that: (1) an encoder network is designed to extract information of geometric structural features and semantic context features of large-scene laser point cloud data, point cloud scene information is fully mined, global description features are generated to form a feature space of point cloud, and similarity measurement of scenes is given according to the distance in the feature space and used for judging whether two or more scene structures are similar or not, and closed-loop detection and relocation of laser SLAM can be realized in the point cloud feature space… (3) the constructed encoder network does not need to be trained in advance according to the point cloud map, and has strong generalization capability”). Thus, it would have been obvious to combine the neural network taught by Mo with the teachings of Hsieh and Miyatani in order to obtain the invention as claimed in Claim 1.
As to Claim 3, Hsieh in view of Miyatani teaches inputting the environmental image data into a neural network model to perform a feature extraction (see Hsieh, paragraph [0040], “Step S22: inputting the plurality of virtual images into a trained deep learning network to perform image feature extractions on the plurality of virtual images”),
Neither Hsieh nor Miyatani fails to teach obtaining the global description information corresponding to the environmental image data, wherein the global description information refers to description information for recording a global feature of the environmental image data. However, Mo teaches inputting image data into a neural network model to perform feature extraction (see paragraph [0085], “the encoder network model and parameters in step 4 are taken, global description features are extracted from the point cloud key frames, NN matching is carried out on the point cloud key frames and the global description features of the historical key frames according to Euclidean distances”),
and obtaining the global description information corresponding to the environmental image data, wherein the global description information refers to description information for recording a global feature of the environmental image data (see paragraph [0008], “The technical problem to be solved by the invention is to provide a point cloud feature space characterization method for laser SLAM, which can extract point cloud global description features representing scene structure information and semantic information”).
Thus, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the network for global description information extraction taught by Mo with the teachings of Hsieh and Miyatani. The motivation for doing so would be to utilize the network in order to extract features from a large set of data (see Mo, paragraph [0040]). Thus, it would have been obvious to combine the neural network taught by Mo with the teachings of Hsieh and Miyatani in order to obtain the invention as claimed in Claim 3.
As to Claim 4, Hsieh teaches each map frame comprises position and orientation information, image feature points, and descriptors of all image feature points of the map frame (see paragraph [0045], “Each virtual image is then used to extract a virtual image feature Fn of the virtual image through VGG, and finally the position and pose Pn={x, y, z, qx, qy, qz, qw} of the virtual camera when each virtual image is generated and is recorded”).
Hsieh and Miyatani fails to explicitly teach wherein a map frame comprises global description information and descriptors of all image feature points of the map frame. However, Mo teaches a map frame can comprise global description information and descriptors of all image feature points (see paragraph [0040], “The invention has the beneficial effects that: (1) an encoder network is designed to extract information of geometric structural features and semantic context features”).
Thus, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the network for global description information extraction taught by Mo with the teachings of Hsieh and Miyatani. The motivation for doing so would be to utilize the network in order to extract features from a large set of data (see paragraph [0040]). Thus, it would have been obvious to combine the neural network taught by Mo with the teachings of Hsieh and Miyatani in order to obtain the invention as claimed in Claim 4.
As to Claim 5, Hsieh teaches wherein the determining K map frames with a maximum similarity to the environmental image data in the visual map according to feature information comprises: calculating a similarity between the environmental image data and each map frame according to the feature information of the environmental image data and feature information of each map frame in the visual map (see paragraph [0049], “Step S52: executing a similarity matching algorithm on the captured image feature and the plurality of virtual image features to obtain a plurality of matching virtual images with relatively high similarities to the captured image from the plurality of virtual images”),
sorting similarities and selecting the K map frames according to the sorted similarities (see paragraph [0050-0051], “The closer S is to 1, the higher the degree of similarity is… After the virtual image with the highest degree of similarity is matched and obtained”).
Hsieh and Miyatani fail to teach calculating a similarity according to global description data. However, Mo teaches that global description information can be used to determine similarity between map frames (see paragraph [0038], “matching of global description features of the point cloud key frames is determined by sequence matching of consecutive N current point cloud frames and consecutive N historical point cloud frames”). Thus, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the global description matching taught by Mo with the teachings of Hsieh and Miyatani. The motivation for doing so would be to use the global features to accurately determine similarity between frames. Mo teaches in paragraph [0040], “The invention has the beneficial effects that: (1) an encoder network is designed to extract information of geometric structural features and semantic context features of large-scene laser point cloud data, point cloud scene information is fully mined, global description features are generated to form a feature space of point cloud, and similarity measurement of scenes is given according to the distance in the feature space and used for judging whether two or more scene structures are similar or not, and closed-loop detection and relocation of laser SLAM can be realized in the point cloud feature space”. Thus, it would have been obvious to combine the global description information taught by Mo with the teachings of Hsieh and Miyatani in order to obtain the invention as claimed in Claim 5.
As to Claim 6, Hsieh teaches wherein the visual map comprises the position and orientation information of each map frame (see paragraph [0045], “Each virtual image is then used to extract a virtual image feature Fn of the virtual image through VGG, and finally the position and pose”).
Hsieh fails to teach determining whether there exists the plurality of similar scenes in the visual map for the current position of the target object according to the K map frames with the maximum similarity comprises: determining whether the K map frames belong to the same scene according to position- orientation distances among the K map frames with the maximum similarity; when the K map frames are determined to belong to a plurality of scenes, determining that there exists the plurality of similar scenes in the visual map for the current position of the target object.
However, Miyatani teaches determining whether there exists the plurality of similar scenes in the visual map for the current position of the target object according to the K map frames with the maximum similarity (see paragraph [0031], “In step S301, a determination is made as to similarity among the plurality of map elements acquired in step S300, thereby determining whether there is a map element including an image captured at a position where wrong position is likely to be estimated in the localization of the imaging apparatus”), comprises
determining whether the K map frames belong to the same scene according to position- orientation distances among the K map frames with the maximum similarity (see paragraph [0032], “The degree of similarity is calculated not for all possible pairs of map elements stored in the storage apparatus 110 but only for particular pairs of map elements between which the distance calculated based on the position/orientation of the camera is equal to or greater than a predetermined threshold value… Conversely, when the distance between map elements is smaller than or equal to the threshold value (or when the distance is smaller than the threshold value), images at these map elements are obtained by capturing a common environment and thus it is natural that the degree of similarity is high”);
when the K map frames are determined to belong to a plurality of scenes determining that there exists the plurality of similar scenes in the visual map for the current position of the target object (see paragraph [0018], “In a case where there are a plurality of map elements each containing an image similar to the image captured at the current position, there is a possibility that a position indicated by a map element different from a map element corresponding to the current position is incorrectly selected as the current position, which results in wrong position estimation in the localization”, and see paragraph [0032], “For example, when the distance between map elements is greater than or equal to the threshold value, if the degree of similarity between images is greater than a threshold value, wrong position is likely to be estimated in the localization. That is, such a situation is undesirable”).
It would have been obvious to combine the plurality of similar scene determination taught by Miyatani with the teachings of Hsieh and Mo. The motivation for doing so would be to prevent incorrect position estimation, as taught by Miyatani in paragraph [0018]. Thus, it would have been obvious to combine the teachings of Miyatani with the teachings of Hsieh and Mo in order to obtain the invention as claimed in Claim 6.
As to Claim 11, Hsieh teaches when there are no symmetrical scene in the K map frames, and there is one map frame among the K map frames which is capable of matching the environmental image data and is capable of being configured to reposition the target object (see paragraph [0051], “After the virtual image with the highest degree of similarity is matched and obtained, in order to solve design issues related to repetitiveness and symmetry of the building, a correct image is manually selected from the plurality of matched virtual images with relatively higher similarities”),
obtaining an initial position of the target object (see paragraph [0089], “Step S103: obtaining the virtual image corresponding to the tracking captured image from the plurality virtual images according to the updated capturing position and the updated capturing pose parameter”).
Hsieh fails to teach determining there is no plurality of scenes. Furthermore, Hsieh fails to teach setting the current position of the target object as a relocatable point, repositioning is performed based on the relocatable point, and obtaining an initial position of the target object. However, Miyatani teaches when there is no plurality of scene map frames in the K map frames (see paragraph [0025], “The determination unit 102 determines whether map elements include images captured at positions where wrong position is likely to be estimated in the localization”), and there is one map frame among the K map frames which is capable of matching the environmental image data and is capable of being configured to reposition the target object setting the current position of the target object as a relocatable point, repositioning is performed based on the relocatable point (see paragraph [0065], “In step S302, based on the information determined in step S301, an image is generated which indicates a candidate for a recovery position when the localization fails”, wherein the ‘candidate’ is the relocatable point).
Thus, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the repositioning taught by Miyatani with the teachings of Hsieh. The motivation for doing so would be to prevent inaccurate localization of a target object (see paragraph [0062]). Thus, it would have been obvious to combine the repositioning taught by Miyatani with the teachings of Hsieh and Mo in order to obtain the invention as claimed in Claim 11.
As to Claim 16, Hsieh fails to explicitly teaches when there exists the plurality of similar scenes or the symmetrical scene in the visual map for the current position of the target object, or the map frame with the maximum similarity among the K map frames is not the target similar scene for the current position, controlling the target object to move, and returning to acquire the environmental image data of the target object to be positioned at the current position.
However, Miyatani teaches when there the map frame with the maximum similarity among the K map frames is not the target similar scene for the current position, controlling the target object to move, and returning to acquire the environmental image data of the target object to be positioned at the current position (see paragraph [0062], “In the localization using the SLAM technology, the localization fails if there is no map element that contains an image partially similar to an image captured at the current position. One method of recovering from this situation is to move the camera to a position corresponding to a map element where it is possible to capture an image which partially similar to another existing image, and restarting the localization at this position”).
Thus, it would have been obvious to combine teachings of Miyatani with the teachings of Hsieh and Mo. The motivation for doing so would be to prevent inaccurate localization of a target object (see paragraph [0062]). Thus, it would have been obvious to combine the repositioning taught by Miyatani with the teachings of Hsieh and Mo in order to obtain the invention as claimed in Claim 16.
As to Claim 17, Hsieh fails to teach further acquiring
M
positions of the target object after
M
consecutive positioning; determining that the target object is repositioned successfully when the
M
consecutive positioning is successful, where
M
is a positive integer; determining that the repositioning of the target object fails when the positioning fails, moving the target object in a preset area, and returning to acquire the environmental image data of the target object to be positioned at the current position.
However, Miyatani teaches moving the target object in a preset area, and returning to acquire the environmental image data of the target object to be positioned at the current position a number of times until the object the object is positioned successfully (see paragraph [0017], “In the apparatus configured to control the movement of the vehicle using the SLAM technology, when recovery from a failure in the localization is performed or when the localization is started, the following are performed repeatedly until it becomes possible to perform the localization: changing the current position of the vehicle, capturing an image of an environment, and searching for a map element including an image similar to the captured image”).
Thus, it would have been obvious to combine the repositioning taught by Miyatani with the teachings of Hsieh and Mo. The motivation for doing so would be to prevent inaccurate localization of a target object (see paragraph [0062]). Thus, it would have been obvious to combine the repositioning taught by Miyatani with the teachings of Hsieh and Mo in order to obtain the invention as claimed in Claim 17.
As to Claim 18, Hsieh teaches a repositioning apparatus (see paragraph [0030], “FIG. 1 is a functional block diagram of an indoor positioning system”), comprising:
a data acquisition module (see Fig. 1, image capturing device 12), configured to acquire environmental image data of a target object to be positioned at a current position (see paragraph [0047], “Step S50: configuring the image capturing device to obtain a captured image at a current position in the target area”, where the image capturing device is interpreted as the ‘target object’);
an information extraction module (see Fig. 1 processor 100), configured to extract feature information from the environmental image data (see paragraph [0051], “After the virtual image with the highest degree of similarity is matched and obtained, in order to solve design issues related to repetitiveness and symmetry of the building, a correct image is manually selected from the plurality of matched virtual images with relatively higher similarities.”);
a map frame determination module (see Fig. 1 processor 100), configured to determine
K
map frames with a maximum similarity to the environmental image data in the visual map according to the global description information, where
K
is an integer, and satisfies
K
≥
1
(see paragraph [0051], “After the virtual image with the highest degree of similarity is matched and obtained, in order to solve design issues related to repetitiveness and symmetry of the building, a correct image is manually selected from the plurality of matched virtual images with relatively higher similarities.”);
a repositioning module (see Fig. 1 processor 100), configured to, when determining that there is no symmetrical scene in the visual map for the current position of the target object according to the
K
map frames with the maximum similarity (see paragraph [0051], “After the virtual image with the highest degree of similarity is matched and obtained, in order to solve design issues related to repetitiveness and symmetry of the building, a correct image is manually selected from the plurality of matched virtual images with relatively higher similarities.”),,
and when a map frame with a maximum similarity among the
K
map frames is determined as a target similar scene for the current position, obtain the position of the target object according to the map frame with the maximum similarity among the
K
map frames (see paragraph [0055], “After the most similar image is successfully matched, the present disclosure can utilize the most similar image to evaluate the position and the pose of the image capturing device 12 when the captured image is taken”, wherein the image capturing device 12 is the ‘target object’).
Hsieh fails to teach that there is no plurality of similar scenes in the visual map for the current position of the target object. Hsieh further fails to explicitly teach repositioning the target object according to the map frame with maximum similarity.
However, in an analogous art, Miyatani teaches a repositioning module (see Fig. 1, determination unit 102) which comprises
determining
K
map frames with a maximum similarity to the environmental image data in the visual map according to the feature information (see paragraph [0019], “a determination is made for each map element of interest as to whether there is another map element including an image similar to an image of the map element of interest”, and see paragraph [0032], “More specifically, feature vectors are extracted from the respective images, and the degree of similarity of the feature vector is calculated and the degree of similarity of the image is defined by the degree similarity of the feature vector”),
and when determining that there is no plurality of similar scenes in the visual map for the current position of the target object according to the K map frames with the maximum similarity (see paragraph [0031], “In step S301, a determination is made as to similarity among the plurality of map elements acquired in step S300, thereby determining whether there is a map element including an image captured at a position where wrong position is likely to be estimated in the localization of the imaging apparatus”),
and when a map frame with a maximum similarity among the K map frames is determined as a target similar scene for the current position, repositioning the target object according to the map frame with the maximum similarity among the K map frames (see paragraph [0062], “In the localization using the SLAM technology, the localization fails if there is no map element that contains an image partially similar to an image captured at the current position. One method of recovering from this situation is to move the camera to a position corresponding to a map element where it is possible to capture an image which partially similar to another existing image, and restarting the localization at this position”)
Thus, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the determination of similar scenes taught by Miyatani with the teachings of Hsieh. The motivation for doing so would be to prevent inaccurate localization of a target object. Miyatani teaches in paragraph [0062], “However, in a case where there are a plurality of map elements including images similar to the image captured at the recovery position, there is a possibility that a map element different from a map element corresponding to the current position is selected, and thus wrong position estimation occurs in the localization at the recover position. To avoid the above situation, it is desirable to select, as the recovery start position, a position which is close to the position of the failure in the localization and which corresponds to a map element to which there is no other map element similar”).
Neither Hsieh or Miyatani explicitly teach extracting global description information from the environmental image data, or determining K map frames with a maximum similarity to the environmental image data in the visual map according to the global description information.
However, in an analogous art, Mo teaches a method for positioning a mobile robot (see paragraph [0086], “performing laser SLAM closed-loop task and repositioning task in the point cloud characteristic space by the point cloud characteristic space characterization method”) which comprises
extracting global description information from the environmental image data (see paragraph [0008], “The technical problem to be solved by the invention is to provide a point cloud feature space characterization method for laser SLAM, which can extract point cloud global description features representing scene structure information and semantic information from complex large-scene laser point clouds”),
or determining K map frames with a maximum similarity to the environmental image data in the visual map according to the global description information (see paragraph [0086], “in the actual SLAM operation, the encoder network model and parameters in step 4 are taken, global description features are extracted from the point cloud key frames, NN matching is carried out on the point cloud key frames and the global description features of the historical key frames according to Euclidean distances”).
Thus, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the network for global description information extraction taught by Mo with the teachings of Hsieh and Miyatani. The motivation for doing so would be to utilize the network in order to extract features from a large set of data (see paragraph [0040], “The invention has the beneficial effects that: (1) an encoder network is designed to extract information of geometric structural features and semantic context features of large-scene laser point cloud data..(3) the constructed encoder network does not need to be trained in advance according to the point cloud map, and has strong generalization capability”). Thus, it would have been obvious to combine the neural network taught by Mo with the teachings of Hsieh and Miyatani in order to obtain the invention as claimed in Claim 18.
As to Claim 19, Hsieh in view of Miyatani and Mo teaches a computer device (see Hsieh, Fig. 1, computing device 10), comprising a processor and a memory storing a computer program (see Hsieh, Fig. 1, processor 100 and storage unit 102), wherein the processor, when executing the computer program implements the method of Claim 1.
As to Claim 20, Hsieh in view of Miyatani and Mo teaches computer-readable storage medium, on which a computer program is stored, wherein the computer program, when executed by a processor (see Hsieh, paragraph [0031], “Furthermore, the storage unit 102 may be, for example, a memory system, which can include a non-volatile memory (such as flash memory) and a system memory (such as DRAM)”), causes the processor to implement the method of Claim 1.
Claim 2 is rejected under 35 U.S.C. 103 as being unpatentable over Hsieh et al. (US Pub No 2022/0301222), hereinafter Hsieh, in view of Miyatani (US Pub No 2023/0168103), hereinafter Miyatani, and further in view of Mo et al. (CN Pub No 111783838), hereinafter Mo, and further in view of Ha et al. (US Pub No 2023/0010105 ), hereinafter Ha.
As to Claim 2, Hsieh in view of Miyatani and Mo teaches that the environmental image data may be an image captured by a camera installed on the target object (see Miyatani, paragraph [0027], “Each map element is a key frame that associates an image of the environment captured by a camera mounted on the vehicle”)
Hsieh in view of Miyatani and Mo fails to teach wherein the environmental image data is a top-view visual image of the target object, and the top-view visual image is captured by a camera installed above the target object. However, in an analogous art, Ha teaches a method for positioning a robot (see paragraph [0001], “ The present disclosure generally relates to the technology of simultaneous localization and mapping (SLAM) in an environment, and in particular, to systems and methods for generating an initial map for a mobile robot with respect to its environment using image data”),
which comprises obtaining image data captured by a camera installed on the target object (see paragraph [0034], “In some embodiments, the mobile robot 102 is equipped with both a front view camera (e.g., forward facing) and a top view camera (upward facing) to capture images at different perspectives in the environment 100”, where the mobile robot is interpreted as the target object.).
Thus, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine top-view camera taught by Ha with the teachings of Hsieh, Miyatani, and Mo. The motivation for doing so would be obtain image data with different perspectives (see paragraph [0034]). Thus, it would have been obvious to combine the top-view camera taught by Ha with the teachings of Hsieh, Miyatani, and Mo in order to obtain the invention as claimed in Claim 2.
Claim 7 is rejected under 35 U.S.C. 103 as being unpatentable over Hsieh et al. (US Pub No 20220301222), hereinafter Hsieh, in view of, Miyatani (US Pub No 2023/0168103), hereinafter Miyatani, further in view of Mo et al. (CN Pub No 111783838), hereinafter Mo, and further in view of Yang et al. (CN Pub No 109141393), hereinafter Yang.
As to Claim 7, Hsieh in view of Miyatani and Mo teaches wherein the determining whether the K map frames belong to the same scene according to the position-orientation distances among the K map frames with the maximum similarity comprises:
calculating a distance between position and orientation information of the map frame with the maximum similarity among the K map frames and position and orientation information of each of other map frames among the K map frames other than the map frame with the maximum similarity, and obtaining K-1 position-orientation distances (see paragraph [0032], “The degree of similarity is calculated not for all possible pairs of map elements stored in the storage apparatus 110 but only for particular pairs of map elements between which the distance calculated based on the position/orientation),
and determining which map frap among the K-1 position orientation distances are greater than a first threshold (see paragraph [0032], “For example, when the distance between map elements is greater than or equal to the threshold value, if the degree of similarity between images is greater than a threshold value, wrong position is likely to be estimated in the localization. That is, such a situation is undesirable.”),
Hsieh in view of Miyatani and Mo fails to teach when a quantity proportion of position-orientation distances among the K-1 position- orientation distances that are greater than a first threshold value is greater than a second threshold value, determining that the K map frames belong to a plurality of scenes; when the quantity proportion of position-orientation distances among the K map frames that are greater than the first threshold value is less than or equal to the second threshold value, determining that the K map frames belong to the same scene.
However, in an analogous art, Yang teaches a method for relocating an object (see abstract, “The invention provides a relocating method, equipment and a storage medium”),
which comprises calculating a distance between position and orientation information between K map frames when a quantity of the K map frames when a quantity proportion of position-orientation distances among the K-1 position- orientation distances that are less than a first threshold value is greater than a second threshold value, determining that the K map frames belong to the same scene (see paragraph [0062], “In this step, when all the key frames include a part of key frames whose distance from the current position of the robot is less than or equal to the preset distance threshold, the relocation may be completed only according to the part of key frames”, where all of the frames is the first threshold value, and the distance threshold is the second threshold, and where repositioning only occurs if the map frames correspond to the same scene);
when a quantity proportion of position-orientation distances among the K-1 position- orientation distances that are less than a first threshold value is greater than a second threshold value, determining that the K map frames belong to a plurality of scenes (see paragraph [0067], “If all the key frames do not include a part of key frames whose distance from the current position of the robot is less than or equal to the preset distance threshold, step 303 is executed”).
Thus, it would have been obvious to modify the first and second threshold and combine the first and second threshold taught by Yang with the teachings of Hsieh, Miyatani, and Mo. The motivation for doing so would be to reduce the calculation time needed to determine the correct frame for positioning. Yang teaches in paragraph [0064], “Here, when there is only one partial key frame, the partial key frame is directly used as a target key frame, and the motion estimation of the current image frame relative to the target key frame is determined according to the target key frame and the current image frame, so that the feature point matching processing of the current image frame and the key frame can be avoided, and the calculation amount and the positioning time consumption are further reduced” Thus, it would have been obvious to combine the teachings of Yang with the teachings of Hsieh, Miyatani, and Mo in order to obtain the invention as claimed in Claim 7.
Claim 12 is rejected under 35 U.S.C. 103 as being unpatentable over Hsieh et al. (US Pub No 2022/0301222), hereinafter Hsieh, in view of, Miyatani (US Pub No 2023/0168103) , hereinafter Miyatani, further in view of Mo et al. (CN Pub No 111783838), hereinafter Mo, and further in view of Guo et al. (US Pub No 2023/0281864), hereinafter Guo.
As to Claim 12, Hsieh in view of Miyatani and Mo fails to teach, wherein the symmetrical scene refers to a map frame with a symmetrical texture among the K map frames.
However, in an analogous art of localization, Guo teaches determining if an object in an image frame comprises a symmetrical texture (see paragraph [0029], “As shown in FIG. 3A, each object in the first row has a texture (e.g., product label displayed on product) that causes the object to be classified as asymmetrical”).
Thus, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the symmetrical texture taught by Guo with the teachings of Hsieh, Miyatani and Mo. The motivation for doing so would be to accurately detect symmetric objects. Guo teaches in paragraph [0002], “However, there a number of challenges with respect to detecting semantic keypoints for textureless or symmetric objects because some of their semantic keypoints may become interchanged. Accordingly, the detection of semantic keypoints for those objects across different frames can be highly inconsistent such that they cannot contribute to valid 6DoF poses under the world coordinate system.” Thus, it would have been obvious to combine the symmetrical texture taught by Guo with the teachings of Hsieh, Miyatani and Mo in order to obtain the invention as claimed in Claim 12.
Allowable Subject Matter
Claims 8-10 and 13-15 are objected to as being dependent upon a rejected base claim, but would be allowable if rewritten in independent form including all of the limitations of the base claim and any intervening claims. The following is a statement of reasons for the indication of allowable subject matter:
As to Claim 8, Hsieh teaches determining whether there exists a symmetrical scene in a visual map (see paragraph [0011], “ In this way, a situation in which a position cannot be determined due to repetitive and symmetrical structures can be solved”). However, Hsieh in view of Miyatani and Mo fails to teach wherein the determining whether there exists a symmetrical scene in the visual map for the current position of the target object comprises: performing forward feature point matching on the map frame with the maximum similarity among the K map frames and the environmental image data of the current position to obtain a forward feature point matching ratio, and performing reverse feature point matching on the map frame with the maximum similarity among the K map frames and the environmental image data of the current position to obtain a reverse feature point matching ratio; when a difference value between the forward feature point matching ratio and the reverse feature point matching ratio is less than a third threshold value, determining the map frame with the maximum similarity as the symmetrical scene for the current position.
The closest prior art found, Afrouzi (US Pat No 10386847) teaches a method for determining a heading of a mobile robot which comprises determining if an image frame is symmetrical by counting illuminated points. However, Afrouzi fails to teach extracting feature points, and the performing forward feature matching, and then reverse feature matching. Furthermore, Afrouzi fails to teach obtaining a forward and reverse ratio.
Seo et al. (US Pub No 2013/0322772) teaches a method for performing feature matching on image sets which comprises rotating an image several times, and performing feature matching on the rotated images. However, Seo fails to explicitly teach determining an image is symmetrical through feature matching. Instead, it is known beforehand if an image is symmetrical.
Claims 9-10 are objected to as allowable by virtue of their dependency on Claim 8.
As to Claim 13, Miyatani teaches determining whether there exists a plurality of similar scenes in the visual map for the current position of the target object according to the K map frames with the maximum similarity. However, Miyatani fails to teach performing forward feature point matching on the map frame with the maximum similarity among the K map frames and the environmental image data of the current position to obtain a forward feature point matching ratio, and performing reverse feature point matching on the map frame with the maximum similarity among the K map frames and the environmental image data of the current position to obtain a reverse feature point matching ratio; when a larger one of the forward feature point matching ratio and the reverse feature point matching ratio is less than a fourth threshold value, determining that there exists a plurality of similar scenes in the visual map for the current position of the target object.
The closest prior art found, Afrouzi (US Pat No 10386847) teaches a method for determining a heading of a mobile robot which comprises determining if an image frame is symmetrical by counting illuminated points. However, Afrouzi fails to teach extracting feature points, and the performing forward feature matching, and then reverse feature matching. Instead, the heading is determined by comparing the amount of points with respect to a centerline.
Yu et al. (CN Pub No 112269386) teaches a method for repositioning a robot in a symmetrical environment which comprises rotating image data. However, Yu fails to explicitly teach a first ratio and fourth ratio.
Claims 14-15 are objected to as allowable by virtue of their dependency on Claim 13.
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure.
Chen et al. (CN Pub No 111311588), teaches a method for repositioning a device which comprises calculating a similarity between a current frame and previous map frames, and screening out frames which may correspond to inaccurate positioning during loop-closure.
Zhu et al. (Zhu F, Zheng S, Wang X, He Y, Gui L, Gong L. Real-Time Efficient Relocation Algorithm Based on Depth Map for Small-Range Textureless 3D Scanning. Sensors (Basel). 2019) teaches a method for repositioning which comprises excluding feature points in symmetrical data.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to SOUMYA THOMAS whose telephone number is (571)272-8639. The examiner can normally be reached M-F 8:30-5:00.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Jennifer Mehmood can be reached at (571) 272-2976. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/S.T./ Examiner, Art Unit 2664
/JENNIFER MEHMOOD/ Supervisory Patent Examiner, Art Unit 2664