DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Claim Status
Claims 1-13 are pending for examination in the application filed 09/23/2025. Claims 1, 3, and 11-12 are currently amended and claim 13 is new.
Priority
Acknowledgement is made of Applicant’s claim for foreign priority under 35 U.S.C. 119 (a)-(d). The certified copy has been received in parent application CN 202211305618.X, filing date: 10/24/2022. Acknowledgement is additionally made of the present application as a national stage entry of PCT/CN2023/141246, international filing date: 12/22/2023.
Information Disclosure Statement
The information disclosure statement (IDS) submitted on 01/03/2025 has been considered by the examiner.
Claim Interpretation
The following is a quotation of 35 U.S.C. 112(f):
(f) Element in Claim for a Combination. – An element in a claim for a combination may be expressed as a means or step for performing a specified function without the recital of structure, material, or acts in support thereof, and such claim shall be construed to cover the corresponding structure, material, or acts described in the specification and equivalents thereof.
The following is a quotation of pre-AIA 35 U.S.C. 112, sixth paragraph:
An element in a claim for a combination may be expressed as a means or step for performing a specified function without the recital of structure, material, or acts in support thereof, and such claim shall be construed to cover the corresponding structure, material, or acts described in the specification and equivalents thereof.
The claims in this application are given their broadest reasonable interpretation using the plain meaning of the claim language in light of the specification as it would be understood by one of ordinary skill in the art. The broadest reasonable interpretation of a claim element (also commonly referred to as a claim limitation) is limited by the description in the specification when 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is invoked.
As explained in MPEP § 2181, subsection I, claim limitations that meet the following three-prong test will be interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph:
(A) the claim limitation uses the term “means” or “step” or a term used as a substitute for “means” that is a generic placeholder (also called a nonce term or a non-structural term having no specific structural meaning) for performing the claimed function;
(B) the term “means” or “step” or the generic placeholder is modified by functional language, typically, but not always linked by the transition word “for” (e.g., “means for”) or another linking word or phrase, such as “configured to” or “so that”; and
(C) the term “means” or “step” or the generic placeholder is not modified by sufficient structure, material, or acts for performing the claimed function.
Use of the word “means” (or “step”) in a claim with functional language creates a rebuttable presumption that the claim limitation is to be treated in accordance with 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph. The presumption that the claim limitation is interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is rebutted when the claim limitation recites sufficient structure, material, or acts to entirely perform the recited function.
Absence of the word “means” (or “step”) in a claim creates a rebuttable presumption that the claim limitation is not to be treated in accordance with 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph. The presumption that the claim limitation is not interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is rebutted when the claim limitation recites function without reciting sufficient structure, material or acts to entirely perform the recited function.
Claim limitations in this application that use the word “means” (or “step”) are being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, except as otherwise indicated in an Office action. Conversely, claim limitations in this application that do not use the word “means” (or “step”) are not being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, except as otherwise indicated in an Office action.
This application includes one or more claim limitations that do not use the word “means,” but are nonetheless being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, because the claim limitation(s) uses a generic placeholder that is coupled with functional language without reciting sufficient structure to perform the recited function and the generic placeholder is not preceded by a structural modifier, as explained in MPEP §2181, subsection I (note that the list of generic placeholders below is not exhaustive, and other generic placeholders may invoke 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph)
A. The Claim Limitation Uses the Term “Means” or “Step” or a Generic Placeholder (A Term That Is Simply A Substitute for “Means”)
With respect to the first prong of this analysis, a claim element that does not include the term “means” or “step” triggers a rebuttable presumption that 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, does not apply. When the claim limitation does not use the term “means,” examiners should determine whether the presumption that 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, paragraph 6 does not apply is overcome. The presumption may be overcome if the claim limitation uses a generic placeholder (a term that is simply a substitute for the term “means”). The following is a list of non-structural generic placeholders that may invoke 35 U.S.C. 112(f) or pre- AIA 35 U.S.C. 112, paragraph 6: “mechanism for,” “module for,” “device for,” “unit for,” “component for,” “element for,” “member for,” “apparatus for,” “machine for,” or “system for.” Welker Bearing Co., v. PHD, Inc., 550 F.3d 1090, 1096, 89 USPQ2d 1289, 1293-94 (Fed. Cir. 2008); Massachusetts Inst. of Tech. v. Abacus Software, 462 F.3d 1344, 1354, 80 USPQ2d 1225, 1228 (Fed. Cir. 2006); Personalized Media,161 F.3d at 704, 48 USPQ2d at 1886–87; Mas- Hamilton Group v. LaGard, Inc., 156 F.3d 1206, 1214-1215, 48 USPQ2d 1010, 1017 (Fed. Cir.1998). This list is not exhaustive, and other generic placeholders may invoke 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, paragraph 6.
Such claim limitations in claim 10 are:
A visual semantic vector-based vehicle guidance system, comprising: a classification module, configured to acquire a road image, and classify pixel points in the road image to obtain pixel point categories;
a set partitioning module, configured to perform point set partitioning according to pixel point positions and categories to obtain a plurality of pixel point sets, each pixel point set consisting of the pixel points with continuous positions and the same category;
a coordinate transformation module, configured to project the pixel points in each pixel point set to a ground coordinate system to obtain three-dimensional coordinate values of the pixel points in each pixel point set;
a vectorization module, configured to determine a semantic coordinate and a direction of a corresponding pixel point set as a semantic vector of the pixel point set according to the three-dimensional coordinate value of each pixel point;
and a guidance module, configured to perform road surface marking positioning according to the semantic vector to guide a vehicle to travel.
[pg. 10-11] In some embodiments, the apparatus provided by an embodiment of the present application may be implemented by software. Fig. 2 shows a visual semantic vector-based vehicle guidance system 455 stored in the memory 450, which may be software in forms of programs and plug-ins, and includes the following software modules: a classification module 4551, a set partitioning module 4552, a coordinate transformation module 4553, a vectorization module 4554, and a guidance module 4555. These modules are logical, so they can be randomly combined or further divided according to the realized functions. Functions of various modules are described hereinafter. In some other embodiments, the system provided by an embodiment of the present application may be implemented by hardware. As an example, the system provided by the embodiment of the present application may be a processor in the form of a hardware decoding processor which is programmed to execute the visual semantic vector-based vehicle guidance method provided by the embodiment of this application. For example, the processor in the form of the hardware decoding processor may use one or more Application Specific Integrated Circuits (ASICs), Digital Signal Processors (DSPs), Programmable Logic Devices (PLDs), Complex Programmable Logic Devices (CPLDs), Field Programmable Gate Arrays (FPGAs) or other electronic elements.
Because this/these claim limitation(s) is/are being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, it/they is/are being interpreted to cover the corresponding structure described in the specification as performing the claimed function, and equivalents thereof.
If applicant does not intend to have this/these limitation(s) interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, applicant may: (1) amend the claim limitation(s) to avoid it/them being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph (e.g., by reciting sufficient structure to perform the claimed function); or (2) present a sufficient showing that the claim limitation(s) recite(s) sufficient structure to perform the claimed function so as to avoid it/them being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph.
Claim Rejections - 35 USC § 112
The following is a quotation of 35 U.S.C. 112(b):
(b) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention.
The following is a quotation of 35 U.S.C. 112 (pre-AIA ), second paragraph:
The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the applicant regards as his invention.
Claims 7-8 are rejected under 35 U.S.C. 112(b) or 35 U.S.C. 112 (pre-AIA ), second paragraph, as being indefinite for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor (or for applications subject to pre-AIA 35 U.S.C. 112, the applicant), regards as the invention.
Claim 7, which depends from claim 5, recites the limitations "wherein after the performing PCA on the covariance matrix to obtain a plurality of feature vectors, the method further comprises: ranking feature values corresponding to all feature vectors from large to small, and comparing the top two feature values; and eliminating a corresponding pixel point set in a case that a difference between the top two feature values is less than a preset difference threshold”. There is insufficient antecedent basis for this limitation in the claim.
Claim 8, which depends from claim 5, recites the limitations “wherein after the determining the direction of the pixel point set according to the feature vector with a maximum feature value, the method further comprises: determining contour line information of the corresponding pixel point set according to the position of each pixel point in each pixel point set; and comparing the contour line information with the direction of the pixel point set, and eliminating the corresponding pixel point set in a case that there is no contour line information parallel to the direction of the pixel point set”. There is insufficient antecedent basis for this limitation in the claim.
Claim Rejections - 35 USC § 101
35 U.S.C. 101 reads as follows:
Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title.
Claims 11-12 are rejected under 35 U.S.C. 101 because the claimed invention is directed to non-statutory subject matter. The claim(s) does/do not fall within at least one of the four categories of patent eligible subject matter because claims 11-12 are drawn to “A computer memory or computer-readable storage medium”.
Per the MPEP 2106.03 Eligibility Step 1: The Four Categories of Statutory Subject Matter [R-07.2022], non-limiting examples of claims that are not directed to any of the statutory categories include:
Products that do not have a physical or tangible form, such as information (often referred to as "data per se”) or a computer program per se (often referred to as "software per se") when claimed as a product without any structural recitations; and
Transitory forms of signal transmission (often referred to as "signals per se"), such as a propagating electrical or electromagnetic signal or carrier wave; and
Subject matter that the statute expressly prohibits from being patented, such as humans per se, which are excluded under The Leahy-Smith America Invents Act (AIA ), Public Law 112-29, sec. 33, 125 Stat.284 (September 16, 2011).
Therefore, since claims 11-12 recite a product that does not have a physical or tangible form, it does not fall within a statutory category. Claims 11-12 are not eligible subject matter under 35 USC § 101. The broadest reasonable interpretation of the claim in light of the specification and Official Gazette Notice (1251 OG 212, made available February 23, 2010), concludes that the claim as a whole covers a transitory signal, which does not fall within the definition of a process, machine, manufacture, or composition of matter (In re Nuijten). In view of the Official Gazette Notice (1251 OG 212, made available February 23, 2010), the examiner suggests amending the claims to recite "a non-transitory computer readable storage medium".
Claim Rejections - 35 USC § 102
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action:
A person shall be entitled to a patent unless –
(a)(1) the claimed invention was patented, described in a printed publication, or in public use, on sale, or otherwise available to the public before the effective filing date of the claimed invention.
(a)(2) the claimed invention was described in a patent issued under section 151, or in an application for patent published or deemed published under section 122(b), in which the patent or application, as the case may be, names another inventor and was effectively filed before the effective filing date of the claimed invention.
Claims 1, 3, and 10-12 are rejected under 35 U.S.C. 102(a)(2) as being anticipated by Liu (US20210150203A1).
Regarding claim 1, Liu teaches a visual semantic vector-based vehicle guidance method performed by a processor, comprising ([0005] According to an aspect of the present invention, a method is provided for producing a road layout model. The method includes capturing a plurality of sequential digital images using a video camera, wherein the digital images are of a perspective view. The method also includes converting each of the plurality of sequential digital images into a top-down view image using a processor. [0152] The output model can be displayed to a user on a screen to allow the user to detect immediate upcoming hazards, make travel decisions, and adjust their driving style to adapt to the road conditions on a real time basis. [0138] In various embodiments, g is a 6-layer convolutional neural network (CNN) that converts a semantic top-view into a 1-dimensional feature vector):
acquiring a road image, and classifying pixel points in the road image to obtain pixel point categories ([0069] At block 215, the captured image(s) can be analyzed using a deep learning-based perception systems that can provide pixel accurate semantic segmentation, where scene parsing can provide an understanding of the captured scene, and predict a label, location, and shape for each element in an image. Semantic segmentation can be accomplished using deep neural networks (e.g., fully convolutional neural networks) to annotate each pixel in an image as belonging to a single class/object, where the semantic segmentation can be applied in/to the perspective view);
performing point set partitioning according to pixel point positions and categories to obtain a plurality of pixel point sets, each pixel point set consisting of the pixel points with continuous positions and the same category ([0069] With semantic segmentation, the objects in the annotated images would not overlap each other. The identified and annotated object can be shaded with a specific color to differentiate that object from other adjacent/adjoining objects. Each object in an image can be segmented by clustering the pixels into their ground truth classes. Conditional random field (CRF) can be used for post processing to refine the segmentation result);
projecting the pixel points in each pixel point set to a ground coordinate system to obtain three-dimensional coordinate values of the pixel points in each pixel point set ([0076] At block 220, a top view representation of a road can be generated. A top-view image of the road can be generated by (1) back-projecting all the road pixels into a 3D point cloud);
determining a semantic coordinate and a direction of a corresponding pixel point set as a semantic vector of the pixel point set according to the three-dimensional coordinate value of each pixel point;
PNG
media_image1.png
462
483
media_image1.png
Greyscale
and performing road surface marking positioning according to the semantic vector to guide a vehicle to travel ([0152] The output model can be displayed to a user on a screen to allow the user to detect immediate upcoming hazards, make travel decisions, and adjust their driving style to adapt to the road conditions on a real time basis).
Regarding claim 3, Liu teaches the method of claim 1. Liu further teaches wherein the performing point set partitioning according to pixel point positions and categories to obtain a plurality of pixel point sets comprises: acquiring all pixel points with the same category and the positions of the pixel points to form an initial set ([0069] At block 215, the captured image(s) can be analyzed using a deep learning-based perception systems that can provide pixel accurate semantic segmentation, where scene parsing can provide an understanding of the captured scene, and predict a label, location, and shape for each element in an image. Semantic segmentation can be accomplished using deep neural networks (e.g., fully convolutional neural networks) to annotate each pixel in an image as belonging to a single class/object, where the semantic segmentation can be applied in/to the perspective view);
and selecting at least one pixel point from the initial set as an initial point, placing pixel points adjacent to the initial point into the same subset, and continuing to perform adjacent pixel point retrieving by taking the pixel points in the subset as basis points to obtain a plurality of subsets, each subset serving as a pixel point set ([0069] The identified and annotated object can be shaded with a specific color to differentiate that object from other adjacent/adjoining objects. Each object in an image can be segmented by clustering the pixels into their ground truth classes. Conditional random field (CRF) can be used for post processing to refine the segmentation result).
Regarding claim 10, Liu teaches a visual semantic vector-based vehicle guidance system, comprising: ([0006] According to another aspect of the present invention, a system is provided for producing a road layout model. The system includes, one or more processor devices; a memory in communication with at least one of the one or more processor devices. [0152] The output model can be displayed to a user on a screen to allow the user to detect immediate upcoming hazards, make travel decisions, and adjust their driving style to adapt to the road conditions on a real time basis. [0138] In various embodiments, g is a 6-layer convolutional neural network (CNN) that converts a semantic top-view into a 1-dimensional feature vector. [0176] In other embodiments, the hardware processor subsystem can include dedicated, specialized circuitry that performs one or more electronic processing functions to achieve a specified result. Such circuitry can include one or more application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), and/or programmable logic arrays (PLAs)):
a classification module, configured to acquire a road image, and classify pixel points in the road image to obtain pixel point categories ([0064] FIG. 2 is a block/flow diagram illustrating a system/method for converting a perspective image of a scene to a top view model, in accordance with an embodiment of the present invention. [0069] At block 215, the captured image(s) can be analyzed using a deep learning-based perception systems that can provide pixel accurate semantic segmentation, where scene parsing can provide an understanding of the captured scene, and predict a label, location, and shape for each element in an image. Semantic segmentation can be accomplished using deep neural networks (e.g., fully convolutional neural networks) to annotate each pixel in an image as belonging to a single class/object, where the semantic segmentation can be applied in/to the perspective view);
a set partitioning module, configured to perform point set partitioning according to pixel point positions and categories to obtain a plurality of pixel point sets, each pixel point set consisting of the pixel points with continuous positions and the same category ([0064] FIG. 2 is a block/flow diagram illustrating a system/method for converting a perspective image of a scene to a top view model, in accordance with an embodiment of the present invention. [0069] Semantic segmentation can be accomplished using deep neural networks (e.g., fully convolutional neural networks) to annotate each pixel in an image as belonging to a single class/object, where the semantic segmentation can be applied in/to the perspective view. With semantic segmentation, the objects in the annotated images would not overlap each other. The identified and annotated object can be shaded with a specific color to differentiate that object from other adjacent/adjoining objects. Each object in an image can be segmented by clustering the pixels into their ground truth classes. Conditional random field (CRF) can be used for post processing to refine the segmentation result);
a coordinate transformation module, configured to project the pixel points in each pixel point set to a ground coordinate system to obtain three-dimensional coordinate values of the pixel points in each pixel point set ([0064] FIG. 2 is a block/flow diagram illustrating a system/method for converting a perspective image of a scene to a top view model, in accordance with an embodiment of the present invention. [0076] At block 220, a top view representation of a road can be generated. A top-view image of the road can be generated by (1) back-projecting all the road pixels into a 3D point cloud);
a vectorization module, configured to determine a semantic coordinate and a direction of a corresponding pixel point set as a semantic vector of the pixel point set according to the three-dimensional coordinate value of each pixel point;
PNG
media_image1.png
462
483
media_image1.png
Greyscale
and a guidance module, configured to perform road surface marking positioning according to the semantic vector to guide a vehicle to travel ([0152] The output model can be displayed to a user on a screen to allow the user to detect immediate upcoming hazards, make travel decisions, and adjust their driving style to adapt to the road conditions on a real time basis. [0006] the processor device is configured to transmit the road layout model to the display screen for presentation to a user).
Regarding claim 12, Liu teaches the method of claim 1. Liu further teaches a computer memory or computer-readable storage medium, having a computer program stored thereon, wherein the computer program implements steps of the visual semantic vector-based vehicle guidance method when executed by a processor ([0007] According to another aspect of the present invention, a non-transitory computer readable storage medium comprising a computer readable program for producing a road layout mode is provided).
Regarding claim 11, Liu teaches the medium of claim 12. Liu further teaches wherein the computer memory is incorporated into a computer device comprising the memory, the processor, and the computer program (Fig. 6. [0170] Embodiments may include a computer program product accessible from a computer-usable or computer-readable medium providing program code for use by or in connection with a computer or any instruction execution system. A computer-usable or computer readable medium may include any apparatus that stores, communicates, propagates, or transports the program for use by or in connection with the instruction execution system, apparatus, or device).
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows:
1. Determining the scope and contents of the prior art.
2. Ascertaining the differences between the prior art and the claims at issue.
3. Resolving the level of ordinary skill in the pertinent art.
4. Considering objective evidence present in the application indicating obviousness or nonobviousness.
This application currently names joint inventors. In considering patentability of the claims the examiner presumes that the subject matter of the various claims was commonly owned as of the effective filing date of the claimed invention(s) absent any evidence to the contrary. Applicant is advised of the obligation under 37 CFR 1.56 to point out the inventor and effective filing dates of each claim that was not commonly owned as of the effective filing date of the later invention in order for the examiner to consider the applicability of 35 U.S.C. 102(b)(2)(C) for any potential 35 U.S.C. 102(a)(2) prior art against the later invention.
Claims 2 and 13 are rejected under 35 U.S.C. 103 as being unpatentable over Liu in view of Xu (CN112802111A).
Regarding claim 2, Liu teaches the method of claim 1. Liu further teaches wherein the classifying pixel points in the road image comprises: classifying the road image through a pre-trained neural network to obtain a pixel point category of each pixel point in the road image ([0069] At block 215, the captured image(s) can be analyzed using a deep learning-based perception systems that can provide pixel accurate semantic segmentation, where scene parsing can provide an understanding of the captured scene, and predict a label, location, and shape for each element in an image. Semantic segmentation can be accomplished using deep neural networks (e.g., fully convolutional neural networks) to annotate each pixel in an image as belonging to a single class/object, where the semantic segmentation can be applied in/to the perspective view. [0100] In one or more embodiments, a neural network can be trained and utilized to directly predict road features from single perspective RGB image(s), and from sequential RGB images from a video. In various embodiments, the neural network design can include (i) a feature extractor to leverage information from both domains, simulated and real semantic top-views from, and (ii) a domain-agnostic classifier of scene parameters);
generating category code of each pixel point category according to a quantity of the pixel point categories; and marking the road image according to the category code to obtain an image of the road image as a semantic image, and performing point set portioning according to the semantic image ([0069] The identified and annotated object can be shaded with a specific color to differentiate that object from other adjacent/adjoining objects. Each object in an image can be segmented by clustering the pixels into their ground truth classes. Conditional random field (CRF) can be used for post processing to refine the segmentation result. [0075] In various embodiments, given the semantic segmentation, a mask of foreground pixels can be defined, where a pixel in the mask is 1 if and only if the segmentation at that pixel belongs to any of the foreground classes. Otherwise, the pixel in the mask is 0. In order to inform the CNN about which pixels have to be in-painted, we apply the mask on the input RGB image and define each pixel in the masked input).
Liu does not explicitly teach marking the road image according to the category code to obtain a gray-scale image of the road image as a semantic image.
Xu, in the same field of endeavor of vehicle image processing, teaches marking the road image according to the category code to obtain a gray-scale image of the road image as a semantic image ([pg. 6 para. 4] according to the semantic three-dimensional reconstruction method constructing the reference three-dimensional map, firstly needs to determine the reference point cloud data of reference object model in the reference three-dimensional map semantic feature; namely firstly performing semantic segmentation to the image data of the reference object model, obtaining the graph semantic segmentation result, as shown in FIG. 4, the pixel of each image data is given a semantic meaning, and using different gray scale to represent the semantic. [pg. 5 para. 2] In the present invention, for automatic driving, semantic feature refers to those that can make the unmanned vehicle better understand the driving rule, sensing road traffic condition, planning the driving route, and covered in the multi-high precision of the map map, rich-dimensional information. Simple say, street lamp, traffic lights can be referred to as a semantic feature).
Therefore, it would have been obvious to a person of ordinary skill in the art before the time of filing to modify the method of Liu with the teachings of Xu to obtain a gray-scale image of the road image as a semantic image because "semantic segmentation is for pixel-level classification of the image so as to distinguish different partitions, specifically, each pixel of the picture is divided into semantic meaning part, then each part is marked as one of predefined categories on the semantics" [pg. 8 para. 1].
Regarding claim 13, Liu and Xu teach the method of claim 2. Xu further teaches wherein the performing point set partitioning according to pixel point positions and categories to obtain a plurality of pixel point sets comprises: acquiring all pixel points with the same category and the positions of the pixel points to form an initial set ([0069] At block 215, the captured image(s) can be analyzed using a deep learning-based perception systems that can provide pixel accurate semantic segmentation, where scene parsing can provide an understanding of the captured scene, and predict a label, location, and shape for each element in an image. Semantic segmentation can be accomplished using deep neural networks (e.g., fully convolutional neural networks) to annotate each pixel in an image as belonging to a single class/object, where the semantic segmentation can be applied in/to the perspective view);
and selecting at least one pixel point from the initial set as an initial point, placing pixel points adjacent to the initial point into the same subset, and continuing to perform adjacent pixel point retrieving by taking the pixel points in the subset as basis points to obtain a plurality of subsets, each subset serving as a pixel point set ([0069] The identified and annotated object can be shaded with a specific color to differentiate that object from other adjacent/adjoining objects. Each object in an image can be segmented by clustering the pixels into their ground truth classes. Conditional random field (CRF) can be used for post processing to refine the segmentation result).
Claim 4 is rejected under 35 U.S.C. 103 as being unpatentable over Liu in view of Pastore (US20220383132A1).
Regarding claim 4, Liu teaches the method of claim 3. Liu does not explicitly teach wherein after the performing point set partitioning according to pixel point positions and categories to obtain a plurality of pixel point sets, the method comprises: acquiring a centroid of each pixel point set, and calculating a distance between every two centroids; and merging corresponding pixel point sets in a case that the distance between the two centroids is less than a preset distance threshold.
Pastore, in the same field of endeavor of semantic learning, teaches wherein after the performing point set partitioning according to pixel point positions and categories to obtain a plurality of pixel point sets, the method comprises: acquiring a centroid of each pixel point set, and calculating a distance between every two centroids ([0052] In a step 210, each computing device, e.g., each of the first, second, and third computers 102a 102b, 102c, runs its sample data through the trained autoencoder to produce autoencoder outputs. The autoencoder outputs may be high-level representations of the input data and may, in particular, be vectors. The vectors may include multiple variables or parameters, e.g., three or more variables or parameters and even up to 100 or more variables or parameters. These variables or parameters may be referred to as feature values. In a deep learning model that identifies pictures of animals, the autoencoder may recognize various features about each image such as size, number of appendages, ear shape, etc. that may help a deep learning model to classify an animal. These features may be the variables or parameters that are determined by feeding the samples through the autoencoder which analyzes the sample. In the embodiment where images are fed to the autoencoder, the autoencoder analyzes the image and can analyze pixels of the image. [0054] In at least some embodiments, each cluster of sample data or of the autoencoder outputs will have a centroid, i.e. a center of the cluster, and will have a radius. [Abstract] The aggregator may integrate the cluster information to define data classes for machine learning classification. The integrating may include computing a respective distance between centroids of the clusters in order to determine a total number of the data classes);
and merging corresponding pixel point sets in a case that the distance between the two centroids is less than a preset distance threshold ([0058] In at least some embodiments, the cluster information may include centroid information that relates to the centroids of the clusters and the aggregator may compare the centroid information to identify any redundant clusters. For example, if centroids from various parties lie within a distance smaller than a pre-defined new cluster threshold distance, then the aggregator may consider the centroids to belong to redundant clusters that should be consolidated or merged for the tally of classes).
Therefore, it would have been obvious to a person of ordinary skill in the art before the time of filing to modify the method of Liu with the teachings of Pastore to acquire centroids of each pixel point set and merge sets that have centroids within a preset distance because "The aggregator may send a deep learning model that includes an output layer that has a total number of nodes equal to the total number of the data classes. The deep learning model may be for the distributed computing devices to perform machine learning classification in federated learning" [0006].
Claim 5 is rejected under 35 U.S.C. 103 as being unpatentable over Liu in view of Li (US20230316742A1).
Regarding claim 5, Liu teaches the method of claim 1. Liu further teaches wherein the projecting the pixel points in each pixel point set to a ground coordinate system to obtain three-dimensional coordinate values of the pixel points in each pixel point set comprises: acquiring an intrinsic matrix of an image collection device for shooting the road image ([0077] In various embodiments, given a single perspective image as input, a corresponding per-pixel semantic segmentation on background classes as well as dense depth map can be obtained. Combined with the intrinsic camera parameters, each coordinate of the perspective view can be mapped into the 3D space);
mapping the position of each pixel point in the pixel point set to a coordinate system of the image collection device according to the intrinsic matrix, and allocating a preset depth value for each pixel point to obtain a pixel point coordinate value in the coordinate system of the image collection device ([0080] In various embodiments, two decoders can be put on fused feature representation(s) for predicting semantic segmentation and the depth map of an occlusion-free scene. Given the depth map and the intrinsic camera parameters, each coordinate of the perspective view can be mapped into the 3D space. The z-coordinate (height axis) can be dropped for each 3D point);
and projecting coordinate values of the pixel points in the coordinate system of the image collection device to the ground coordinate system to obtain the three-dimensional coordinate value of each pixel point in the pixel point set ([0076] At block 220, a top view representation of a road can be generated. A top-view image of the road can be generated by (1) back-projecting all the road pixels into a 3D point cloud).
Liu does not explicitly teach acquiring an intrinsic matrix and an extrinsic matrix of an image collection device for shooting the road image; and projecting coordinate values of the pixel points in the coordinate system of the image collection device to the ground coordinate system according to the extrinsic matrix to obtain the three-dimensional coordinate value of each pixel point in the pixel point set.
Li, in the same field of endeavor of vehicle image processing, teaches acquiring an intrinsic matrix and an extrinsic matrix of an image collection device for shooting the road image ([0102] The camera parameter matrix comprises an intrinsic matrix and an extrinsic matrix, wherein the intrinsic matrix is related to the camera and comprises a focal length, a principal point coordinate position relative to an imaging plane, a coordinate axis inclination parameter and a distortion parameter, the intrinsic matrix reflects the attributes of the cameras, and the intrinsic matrices of the cameras are different. The extrinsic matrix depends on a position of the camera in the world coordinate system and includes a rotation matrix and a translation matrix, which together describe how to convert points from the world coordinate system to the camera coordinate system. [0003] for example, lane line detection, road segmentation, target detection and track planning, so as to obtain perception task results, wherein the perception task results are of great significance in fields of traffic and the like, and the driving safety of the autonomous vehicle can be ensured);
and projecting coordinate values of the pixel points in the coordinate system of the image collection device to the ground coordinate system according to the extrinsic matrix to obtain the three-dimensional coordinate value of each pixel point in the pixel point set ([0153] In some embodiments, the S301 described above may further be implemented by the following: obtaining an intrinsic matrix and an extrinsic matrix of the acquisition apparatus: performing feature extraction on the plurality of view images to obtain the plurality of view image features; converting each two-dimensional coordinate point in the preset aerial view query vector into a plurality of three-dimensional spatial points to obtain a plurality of three-dimensional spatial points corresponding to two-dimensional coordinate points; projecting the plurality of three-dimensional spatial points corresponding to the two-dimensional coordinate points to at least one view image according to the intrinsic matrix and the extrinsic matrix to obtain a plurality of projection points of the two-dimensional coordinate points; and performing dense sampling in the plurality of view image features and performing feature fusion thereon to obtain the second aerial view query vector according to the plurality of projection points of the two-dimensional coordinate points. [Abstract] wherein the preset aerial view query vector corresponds to a three-dimensional physical world which is a preset range away from the vehicle in a real scene at the current moment).
Therefore, it would have been obvious to a person of ordinary skill in the art before the time of filing to modify the method of Liu with the teachings of Li to acquire an intrinsic and extrinsic matrix of an image collection device and project coordinate values of the pixel points to the ground coordinate system according to the extrinsic matrix to obtain the three-dimensional coordinate value so that "then projection points can be obtained on the plurality of view images in a projection mode, that is, the two-dimensional position points and the plurality of view images have corresponding relations, and an aerial view query vector with temporal information and spatial information fused can be obtained by performing sampling and interactive fusing on the projection points, and is represented by the aerial view query vector at the current moment" [0060].
Claim 6 is rejected under 35 U.S.C. 103 as being unpatentable over Liu in view of Bastos (US20230072966A1).
Regarding claim 6, Liu teaches the method of claim 1. Liu does not explicitly teach wherein the determining a semantic coordinate and a direction of a corresponding pixel point set as a semantic vector of the pixel point set according to the three-dimensional coordinate value of each pixel point comprises: determining a centroid of the pixel point set according to the three-dimensional coordinate values of the pixel points in the pixel point set; determining a covariance matrix of the pixel point set according to an offset between each pixel point and the centroid in the pixel point set; performing Principal Component Analysis (PCA) on the covariance matrix to obtain a plurality of feature vectors; and determining the direction of the pixel point set according to the feature vector with a maximum feature value, and determining the semantic vector of the pixel point set in combination with the direction of the pixel point set by taking a coordinate of the centroid as a semantic coordinate.
Batsos, in the same field of endeavor of semantic labeling, teaches wherein the determining a semantic coordinate and a direction of a corresponding pixel point set as a semantic vector of the pixel point set according to the three-dimensional coordinate value of each pixel point comprises: determining a centroid of the pixel point set according to the three-dimensional coordinate values of the pixel points in the pixel point set ([0027] The robotic systems may use the machine learning models/algorithms to facilitate (i) semantic label predictions for LiDAR point clouds and (ii) the generation of a 3D representation of a scene over multiple LiDAR sweeps. [0067] In 718, the feature vectors F.sub.i are input into an unsupervised machine learning model/algorithm. The unsupervised machine learning algorithm is generally configured to generate a second confidence score S.sub.2 for each pair of data points in the 3D point cloud(s) and semantic label. The unsupervised machine learning model/algorithm can include, but is not limited to, a Principal Component Analysis (PCA) model/algorithm. The PCA model/algorithm is configured to calculate a projection of the feature vector data into a single dimension. This computation is achieved by: calculating the mean value M of each column in a feature vector F.sub.i (i.e., M=mean(F.sub.i)); generate a centered matrix C by subtracting the mean values from the values of the feature vector (i.e., C=F.sub.i−M); calculate a covariance matrix V of the centered matrix (i.e., V=.sub.cov(C)));
determining a covariance matrix of the pixel point set according to an offset between each pixel point and the centroid in the pixel point set; performing Principal Component Analysis (PCA) on the covariance matrix to obtain a plurality of feature vectors ([0067] The PCA model/algorithm is configured to calculate a projection of the feature vector data into a single dimension. This computation is achieved by: calculating the mean value M of each column in a feature vector F.sub.i (i.e., M=mean(F.sub.i)); generate a centered matrix C by subtracting the mean values from the values of the feature vector (i.e., C=F.sub.i−M); calculate a covariance matrix V of the centered matrix (i.e., V=.sub.cov(C)); and calculate the eigen-decomposition of the covariance matrix to obtain eigenvalues and eigenvectors (i.e., values, vectors=eig(V)));
and determining the direction of the pixel point set according to the feature vector with a maximum feature value, and determining the semantic vector of the pixel point set in combination with the direction of the pixel point set by taking a coordinate of the centroid as a semantic coordinate ([0067] The second confidence scores are computed by projecting the feature vector of each data point to the eigenvector with the largest eigenvalue. FIG. 8 shows a line 802 illustrating second confidence scores plotted on a graph. [0006] The unsupervised machine learning algorithm may (i) perform a principal component analysis to identify the most significant eigenvector of the feature vector and (ii) set the second confidence score equal to the projection of the feature vector on the first principle component identified by the eigenvalues).
Therefore, it would have been obvious to a person of ordinary skill in the art before the time of filing to modify the method of Liu with the teachings of Batsos to determine the semantic vector by performing PCA on the covariance matrix to "to identify the most significant eigenvector of the feature vector" [0006] because "the projected data points and semantic labels are analyzed (e.g., manually) to determine whether the projected data points are correctly associated with the respective semantic labels. This analysis can involve: identifying features in the sensor data; and confirming that the semantic labels associated with the features are accurate" [0055].
Claim 9 is rejected under 35 U.S.C. 103 as being unpatentable over Liu in view of Dudley (US20210191394A1).
Regarding claim 9, Liu teaches the method of claim 1. Liu does not explicitly teach wherein after the performing road surface marking positioning according to the semantic vector, the method further comprises: generating a voice call instruction according to the direction of the semantic vector; and outputting corresponding voice information from a preset voice library in response to the voice call instruction to guide the vehicle to travel.
Dudley, in the same field of endeavor of vehicle guidance, teaches wherein after the performing road surface marking positioning according to the semantic vector, the method further comprises: generating a voice call instruction according to the direction of the semantic vector; and outputting corresponding voice information from a preset voice library in response to the voice call instruction to guide the vehicle to travel ([0064] Still further, the derived representation of the surrounding environment perceived by AV 300 may be embodied in various forms. For instance, as one possibility, the derived representation of the surrounding environment perceived by AV 300 may be embodied in the form of a data structure that represents the surrounding environment perceived by AV 300, which may comprise respective data arrays (e.g., vectors) that contain information about the objects detected in the surrounding environment perceived by AV 300, a data array that contains information about AV 300, and/or one or more data arrays that contain other semantic information about the surrounding environment. [0041] At this third time shown in FIGS. 2C-D, AV 200 is still presenting the planned trajectory for AV 200 via the HUD system, which is again displayed as a path extending from the front of AV 200. Additionally, at this third time shown in FIGS. 2C-D, AV 200 performs yet another evaluation of the current scenario being faced by AV 200 in order to determine whether to selectively present any scenario-based information to the local safety driver of the AV, which may again involve an evaluation of factors such as a type of scenario being faced by AV 200, a likelihood of making physical contact with the other vehicles in the AV's surrounding environment in the near future, an urgency level associated with the current scenario, and/or a likelihood that the local safety driver is going to disengage the autonomy system in the near future…Thus, at the third time shown in FIGS. 2C-D, AV 200 is now presenting another curated set of scenario-based information to the local safety driver, which comprises both visual information output via the AV's HUD system that includes a bounding box for stop sign 202, a bounding box and predicted future trajectory for vehicle 203, a bounding box and predicted future trajectory for pedestrian 204, and a stop wall 205 that indicates where AV 200 plans to stop for the stop sign, as well as audio information output via the AV's speaker system notifying the local safety driver that AV 200 has detected an “approaching a stop-sign intersection” type of scenario. [0152] At block 403, in response to determining that the current scenario warrants presentation of scenario-based information to the safety driver of AV 300′, on-board computing system 302 may select a particular set of scenario-based information (e.g., visual and/or audio information) to present to the safety driver of AV 300′).
Therefore, it would have been obvious to a person of ordinary skill in the art before the time of filing to modify the method of Liu with the teachings of Dudley to generating a voice call instruction according to the direction of the semantic vector and output corresponding voice information to guide the vehicle to travel to "advantageously enable a safety driver (or the like) to monitor the status of the AV's autonomy system—which may help the safety driver of the AV make a timely and accurate decision as to whether to switch the AV from an autonomous mode to a manual mode—while at the same time minimizing the risk of overwhelming and/or distracting the safety driver with extraneous information that is not particularly relevant to the safety driver's task" [0171].
Allowable Subject Matter
Claims 7 and 8 are objected to as being dependent upon a rejected base claim but would be allowable if rewritten in independent form including all of the limitations of the base claim and any intervening claims and if all withstanding rejections and objections were overcome.
Regarding claim 7, Bastos teaches performing PCA on the covariance matrix to obtain a plurality of feature vectors ([0067] The PCA model/algorithm is configured to calculate a projection of the feature vector data into a single dimension. This computation is achieved by: calculating the mean value M of each column in a feature vector F.sub.i (i.e., M=mean(F.sub.i)); generate a centered matrix C by subtracting the mean values from the values of the feature vector (i.e., C=F.sub.i−M); calculate a covariance matrix V of the centered matrix (i.e., V=.sub.cov(C)); and calculate the eigen-decomposition of the covariance matrix to obtain eigenvalues and eigenvectors (i.e., values, vectors=eig(V))). The following limitations were not found to be taught in the art: wherein after the performing PCA on the covariance matrix to obtain a plurality of feature vectors, the method further comprises: ranking feature values corresponding to all feature vectors from large to small, and comparing the top two feature values; and eliminating a corresponding pixel point set in a case that a difference between the top two feature values is less than a preset difference threshold.
Regarding claim 8, Bastos teaches determining the direction of the pixel point set according to the feature vector with a maximum feature value ([0067] The second confidence scores are computed by projecting the feature vector of each data point to the eigenvector with the largest eigenvalue. FIG. 8 shows a line 802 illustrating second confidence scores plotted on a graph. [0006] The unsupervised machine learning algorithm may (i) perform a principal component analysis to identify the most significant eigenvector of the feature vector and (ii) set the second confidence score equal to the projection of the feature vector on the first principle component identified by the eigenvalues). The following limitations were not found to be taught in the art: wherein after the determining the direction of the pixel point set according to the feature vector with a maximum feature value, the method further comprises: determining contour line information of the corresponding pixel point set according to the position of each pixel point in each pixel point set; and comparing the contour line information with the direction of the pixel point set, and eliminating the corresponding pixel point set in a case that there is no contour line information parallel to the direction of the pixel point set.
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure.
Agia (US20230267615A1) teaches road surface semantic segmentation.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to Jacqueline R Zak whose telephone number is (571)272-4077. The examiner can normally be reached M-F 9-5.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Emily Terrell can be reached at (571) 270-3717. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/JACQUELINE R ZAK/Examiner, Art Unit 2666
/EMILY C TERRELL/Supervisory Patent Examiner, Art Unit 2666