DETAILED ACTION
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Examiner Notes
Examiner cites particular columns and line numbers in the references as applied to the claims below for the convenience of the applicant. Although the specified citations are representative of the teachings in the art and are applied to the specific limitations within the individual claim, other passages and figures may apply as well. It is respectfully requested that, in preparing responses, the applicant fully consider the references in entirety as potentially teaching all or part of the claimed invention, as well as the context of the passage as taught by the prior art or disclosed by the examiner.
The examiner encourages Applicant to submit an authorization to communicate with the examiner via the Internet by making the following statement (from MPEP 502.03):
“Recognizing that Internet communications are not secure, I hereby authorize the USPTO to communicate with the undersigned and practitioners in accordance with 37 CFR 1.33 and 37 CFR 1.34 concerning any subject matter of this application by video conferencing, instant messaging, or electronic mail. I understand that a copy of these communications will be made of record in the application file.”
Please note that the above statement can only be submitted via Central Fax, Regular postal mail, or EFS Web (PTO/SB/439).
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows:
1. Determining the scope and contents of the prior art.
2. Ascertaining the differences between the prior art and the claims at issue.
3. Resolving the level of ordinary skill in the pertinent art.
4. Considering objective evidence present in the application indicating obviousness or nonobviousness.
Claims 1, 6-8, 13-15, and 20 are rejected under 35 U.S.C. 103 as being unpatentable over Miller et al. (US 20230130770) in view of Michalopulos et al. (US 20210366256).
As per claim 1, Michalopulos teaches the invention substantially as claimed including a system comprising:
a memory configured to store a training dataset comprising a first image of a first physical device ([0009], computing system can utilize the three-dimensional map, eye gazing technique in combination with utilizing the machine learning model to enable the user to access dynamic information including physical, context, and semantic information of all the physical objects or devices in addition to issuing commands for interaction; [0039], computing system 114 may also comprise one or more data stores, for example, object and device library 124 related to storing information and data of physical objects and devices of all the physical environments associated with the user(s); and [0046], each of one or more devices 110a, . . . , 110n, and all the other one or more objects are assigned with confidence distribution score... the confidence distribution score associated with each physical object and object representations in 3D maps may be used in combination with various machine learning models. For example, the confidence distribution score may be one of the attributes or parameters for training and updating machine learning models), wherein the first image is labeled with the first physical device ([0045], all the physical objects including all the one or more devices 110a, . . . , 110n such as smart devices or IoT devices... are associated with one or more tags, info-label data, object types, classifiers, and identifications... all the one or more devices 110a, . . . , 110n, and all the other one or more objects are categorized into one or more categories and stored in data stores or in object and device library 124 of the computing system 114; and [0046], each of one or more devices 110a, . . . , 110n, and all the other one or more objects are assigned with confidence distribution score);
a camera configured to capture images of objects within a field of view of the camera ([0033], image sensors, for example, cameras may be head worn as shown in FIG. 1B... The image sensors are configured to capture images or live streams of the physical environment where the user is around and present. In an embodiment, the image sensors capture the images and/or live streams of the physical environment in real-time and provide the user with viewing 3D object-centric map representation of all physical objects as virtual objects or mixed reality objects on the display 108 of the AR device 102 ; and [0055], HMD 102 may include one or more cameras that can capture images and videos of environments); and
a processor associated with a spatial computing device, operably coupled to the memory and camera ([0027], a distributed system 100 comprises the artificial reality (AR) device 102, and the one or more devices 110a-110n... the one or more processors of the system 100 may be configured to implement each and every functionality of the artificial reality (AR) device 102, any of the devices 110a, 110b, . . . , 110n, and the computing system 114; [0033], the AR device 102 may comprise one or more processors 104, a memory 106 for storing computer-readable instructions to be executed on the one or more processors 104, and a display 108...The AR device 102 may also comprise one or more sensors...the one or more sensors may include but are not limited to, image sensors,...The image sensors, for example, cameras may be head worn as shown in FIG. 1B; [0038], the computing system 114 may be among the one or more devices 110a, . . . 110n and/or maybe a standalone host computer system, an on-board computer system integrated with the AR device 102, graphical user interface, or any other hardware platform capable of providing artificial reality 3D object-centric map representation to and receiving commands associated with the intent of interactions from the user(s) via eye tracking and eye gazing of the user in real-time and dynamically; and [0039], computing system 114 comprises a processor 116 and a memory 118. The processor 116 is programmed to implement computer-executable instructions that are stored in memory 118 or memory units; Examiner Note: Miller’s distributed system 100 provides the functionality of a spatial computing device including enabling three-dimensional human-computer interaction: [0027], one or more processors of the system 100 may be configured to perform and implement any of the functions and operations relating to representing the physical environment in a form of the virtual environment, carrying out one or more techniques including, but not limited to, simultaneous localizing and mapping (SLAM) techniques, eye tracking and eye gazing techniques, devices or objects tracking operations, multi degrees of freedom (DOF) detecting techniques to determine a pose of the device or a gaze of a user, representing three-dimensional (3D) object-centric map of an environment, etc., for enabling the user in the environment to access and interact with any of the one or more devices 110 (110a, 110b, . . . , 100n) and/or one or more objects different from the one or more devices 110, existing in that environment), and configured to:
receive a second image from the camera, wherein the second image shows a second physical device ([0032], the image sensors capture the images and/or live streams of the physical environment in real-time and provide the user with viewing 3D object-centric map representation of all physical objects as virtual objects or mixed reality objects on the display 108 of the AR device 102; and [0042], as the user moves throughout different spaces or zones or regions, artificial reality device 102 must provide synchronized, continuous, and updated feature maps with low latency in order to provide a high-quality, immersive, and enjoyable experience for users);
extract a first set of features from the first image, wherein the first set of features indicates physical attributes of the first physical device ([0047], one or more devices 110a, . . . , 110n and the one or more objects are identifiable when data points and features including context and semantic features of the objects and devices match with predetermined data points and features including context and semantic features of the objects and devices in the object and device library 124 associated with data repository of objects and device 126);
extract a second set of features from the second image, wherein the second set of features indicates physical attributes of the second physical device ([0043], Each physical object in the physical environment is associated with some features and in particular embodiments, the map generation engine 122 may evaluate those features, including but not limited to physical, context, and semantic features, surface attributes, and data points associated with each physical object and other related features of the physical objects present or exist in the physical environment);
compare each of the first set of features with a counterpart feature from among the second set of features ([0047], one or more devices 110a, . . . , 110n and the one or more objects are identifiable when data points and features including context and semantic features of the objects and devices match with predetermined data points and features including context and semantic features of the objects and devices in the object and device library 124 associated with data repository of objects and device 126);
determine, based at least in part upon the comparison, that the first physical device corresponds to the second physical device ([0047], one or more devices 110a, . . . , 110n and the one or more objects are identifiable when data points and features including context and semantic features of the objects and devices match with predetermined data points and features including context and semantic features of the objects and devices in the object and device library 124 associated with data repository of objects and device 126);
generate a first virtual device in a virtual environment, wherein the first virtual device is a virtual representation of the second physical device ([0004], each object is represented as a “virtual object” in the 3D map... the viewer is enabled to interact with each real object with the same experience as they experience during interacting within the real-world environment);
receive a user input that indicates the first virtual device is requested to perform a first operation ([0029], the data and information exchange includes, but is not limited to...the one or more instructions of the users and the one or more commands with the intent of accessing and interaction by the users, between the various user computers and systems; and [0048], the intent identification engine 130 is programmed or configured to determine the intent of the user to access and interact with a particular physical device represented as object representations (e.g., virtual device or objects) in the 3D map. The intent of the user is determined by determining the instructions and one or more contextual signals. The instructions may be explicit (e.g., “turn on a light”) or implicit (e.g., “where did I buy this”));
establish a communication path between the spatial computing device and the second physical device ([0028], the system 100 includes links 134 and a data communication network 112 enabling communication and interoperation of each of the artificial reality (AR) device 102, the one or more devices 110a, . . . , 110n, and the computing system 114 with one another enabling the users to access and interaction with any of the one or more devices 110a, . . . , 110n and the one or more objects in the physical environment; and [0029], the system 100 provides the users to communicate and interact with each of the AR device 102, the one or more devices 110a, . . . , 110n, and the computing system 114 for providing access, instructions, and commands to or from any of the AR device 102, the one or more devices 110a, . . . , 110, the one or more objects and the computing system 114 through an application programming interfaces (API) or other communication channels); and
in response to receiving the user input, communicate a first control signal to the second physical device ([0030], The instructions and the commands of the users for accessing, interacting, and/or operating various devices or objects of the physical environment through the access of a three-dimensional (3D) object-centric map via the AR device 102 may be executed based on the device-specific application protocols and attributes configured for the corresponding devices), wherein the first control signal causes the second physical device to perform the first operation ([0058], The API platform 200 exposes the current eye vector from real-time eye tracking. The API platform 200 exposes the current object that the user is looking at or intending to interact with where such current object is detected through proximity and eye gaze. The API platform 200 issues a command to an object and/or a device in the physical environment to change states for certain objects and/or devices whose states are user changeable, e.g. a smart light bulb).
Miller fails to specifically teach, determine an identity of the second physical device in response to determining that the first physical device corresponds to the second physical device.
However, Michalopulos teaches, determine an identity of the second physical device in response to determining that the first physical device corresponds to the second physical device ([0118], the captured images of the equipment can be compared to a database of images of tubular equipment to identify images that match the captured images).
Miller and Michalopulos are analogous because they are both related to object detection and object virtualization. Miller teaches a method of capturing object images, comparing a captured image to a database of images, generating a virtual device based on the comparison, and controlling a physical device by interacting with the virtual device:
[0008], after localizing the user in the physical environment in terms of the different poses of the AR device, one or more physical objects, for example, smart objects with smart features, including but not limited to, smart television, smart light bulb, LEDs, and other particular physical objects other than the smart devices like chairs, tables, couches, side tables, etc., are detected. Each physical object existing in the physical environment is identified by evaluating various physical and semantic features including surface attributes and data points of each physical object. The evaluation of the various physical, and semantic features, surface attributes, and data points associated with each physical object, an object representation is generated for each of those physical objects in the three-dimensional map; and [0010], computing system can utilize the three-dimensional map, eye gazing technique in combination with utilizing the machine learning model to enable the user to access dynamic information including physical, context, and semantic information of all the physical objects or devices in addition to issuing commands for interaction. This way, the computing system enhances the AR experience for the user via the usage of the AR device and the computing system via the networking environment.
Michalopulos teaches a method of capturing object images, comparing a captured image to a database of images, identifying an image based on the comparison, and generating a virtual device based on an identified physical object:
[0119], artificial neural networks (ANN) models can be used to identify the tubular equipment, where the ANN is trained using the database of images of tubular equipment to correctly identify the equipment from the picture. Examples of ANN architectures to perform this feature can include convolutional neural networks, residual networks, and other similar architectures. The equipment detection ANN can be applied to a portion of the captured image containing the tubular equipment, or to the image as a whole; [0120], measurements can be taken of the tubular equipment by the computer vision system. This feature can be accomplished by applying edge detection or other object detection techniques to locate relevant features of the tubular equipment, measuring the distance in the image between relevant features, and then converting that image distance to a. real-world distance (e.g. through stereo vision to measure distance to the object, by a known distance from the camera to the object, or other similar techniques). Non-limiting examples of features capable of being measured by embodiments include external diameter, length, length of threading between tubular equipment, thread pitch/kind, etc. Other measurements may be relevant for certain types of tubular equipment, such as number of stabilizers, length of certain subsections of the tubular equipment, etc.; and [0164], computer vision system with a processor and memory may receive the visual data from the subsets 1441, 1442 and analyze the visual data to determine and identify the particular detectable characteristics 1447 of a pipe 1444.
It would have been obvious to one having ordinary skill in the art before the effective filing date of the claimed invention that based on the combination, the teachings of Miller would be modified with the object identification mechanism taught by Michalopulos resulting a in a system that can identify, virtualize, and control objects in an environment. Therefore, it would have been obvious to combine the teachings of Miller and Michalopulos.
As per claim 6, Michalopulos teaches, wherein the processor is further configured to:
receive Light Detection and Ranging (LiDAR) data from a LiDAR sensor circuit ([0112], one or more cameras in accordance with the present disclosure can include cameras that are capable of recording distance or ranging information, such as time-of-flight cameras or LIDAR sensors), wherein the LiDAR data indicates a distance of the second physical device to the spatial computing device ([0112], one or more cameras in accordance with the present disclosure can include cameras that are capable of recording distance or ranging information, such as time-of-flight cameras or LIDAR sensors. Such one or more time-of-flight or LIDAR sensors or cameras can be used to provide accurate distance, size, shape, dimensions, and other important physical information about a person or thing); and
determine the distance between the second physical device to the spatial computing device based at least in part upon the LiDAR data ([0112], one or more cameras in accordance with the present disclosure can include cameras that are capable of recording distance or ranging information, such as time-of-flight cameras or LIDAR sensors. Such one or more time-of-flight or LIDAR sensors or cameras can be used to provide accurate distance, size, shape, dimensions, and other important physical information about a person or thing).
As per claim 7, Miller teaches, wherein the first operation comprises communicating a particular signal to a server ([0031], A user at the AR device 102, and/or the one or more devices 110a, . . . , 110n, and/or the computing system 114 may enter a Uniform Resource Locator (URL) or other address directing the web browser to a particular server...The server may accept the HTTP request and communicate to each of the AR devices 102, the one or more devices 110a, . . . , 110n, the computing system 114 and the system 100, one or more Hyper Text Markup Language (HTML) files responsive to the HTTP request. A webpage is rendered based on the HTML files from the server for presentation to the user with the intent of providing access and issuing commands for interaction with the one or more devices 110, . . . , 110n including the AR device 102, and the computing system 114; and [0058], API platform 200 issues a command to an object and/or a device in the physical environment to change states for certain objects and/or devices whose states are user changeable, e.g. a smart light bulb).
As per claim 8, this is the “method claim” corresponding to claim 1 and is rejected for the same reasons. The same motivation used in the rejection of claim1 is applicable to the instant claim.
As per claim 13, this claim is similar to claim 6 and is rejected for the same reasons.
As per claim 14, this claim is similar to claim 7 and is rejected for the same reasons.
As per claim 15, this is the “non-transitory computer-readable medium claim” corresponding to claim 1 and is rejected for the same reasons. The same motivation used in the rejection of claim1 is applicable to the instant claim.
As per claim 20, this claim is similar to claim 6 and is rejected for the same reasons.
Claims 2-5, 9-12, and 16-19 are rejected under 35 U.S.C. 103 as being unpatentable over Miller-Michalopulos as applied to independent claims 1, 8, and 15 and in further view of Jeong et al. (US 20240104882).
As per claim 2, Miller teaches, wherein the processor is further configured to:
display the first virtual device on a graphical user interface such that a location of the first virtual device in the graphical user interface is tethered to a location of the second physical device in the global plane ([0033], the image sensors capture the images and/or live streams of the physical environment in real-time and provide the user with viewing 3D object-centric map representation of all physical objects as virtual objects or mixed reality objects on the display 108 of the AR device 102 as shown in FIG. 1A; and [0088], Upon clicking, the tablet becomes a 3D map generator (or Inspector) for generating the 3D object-centric map. The 3D map generator surfaces the attributes and live state for any physical object within the room and displays an overview or an inspector view showing the name and image of the object in addition to dynamic information to reflect the location of that object, the time the object was last seen, and the distance between the user and the object. By looking at the objects and clicking on the EMG wireless clicker, any object's information is viewable and accessed).
The combination of Miller-Michalopulos fails to specifically teach, determine a first pixel location coordinate associated with the second physical device in the first image; and determine a first physical location coordinate of the second physical device in a global plane based at least in part upon the first pixel location coordinate associated with the second physical device.
However, Jeong teaches, determine a first pixel location coordinate associated with the second physical device in the first image ([0011], comparing the feature vectors included in the feature map with reference feature vectors generated by the feature generation model based on reference points within a reference image, wherein the reference image includes an reference object instance that corresponds to the target object; based on the comparing, identifying points of interest in the input image that correspond to the reference points; and determining a presence of the target object in the environment based on the comparing; and [0045], Each image frame (also referred to as an “image” herein) can be organized as two-dimensional (2D) image data arranged as an X by Y array of picture element data structures (e.g., pixels) that are each indexed by a respective (x,y) coordinate pair, where each pixel stores image data that represents one or move values about light reflected from a corresponding point in an observed scene... pixel-level image data can take the form of a vector of elements that each represent a respective dimension or channel); and
determine a first physical location coordinate of the second physical device in a global plane based at least in part upon the first pixel location coordinate associated with the second physical device ([0072], corresponding point coordinates for input image 302 can be mapped to a real-time physical location within the environment 102, and coordinates for the physical location can be included in the PPI list 310 or determined at a later stage from the data included in the POI list 310).
The combination of Miller-Michalopulos and Jeong are analogous because they are each related to object detection and object virtualization. Miller teaches a method of capturing object images, comparing a captured image to a database of images, generating a virtual device based on the comparison, and controlling a physical device by interacting with the virtual device. Michalopulos teaches a method of capturing object images, comparing a captured image to a database of images, identifying an image based on the comparison, and generating a virtual device based on an identified physical object. Jeong teaches a method of identifying and controlling objects in an environment including pixel related features:
[0023], performing a physical action in respect of the target object based on the comparing; [0040], the system 100 can enable target objects to be detected, recognized and tracked based on exposure to a reference image of a same or similar object; and [0116], images (for examples images collected by one or more cameras 808(i) within the environment 102) are provided to the correspondence module 110 where each image can be processed as follows: trained feature generation model 124 can be used to generate a feature map 306 of pixel-level feature vectors for the image; the feature map 306 is then searched using corresponding point detection operation 126 to determine pixel locations having feature vectors that match reference point feature vectors in the reference point feature vector list 210. Identified matches are then output as points-of-interest in a point-of-interest list 310 generated for the image.
It would have been obvious to one having ordinary skill in the art before the effective filing date of the claimed invention that based on the combination, the teachings of the combination of Miller- Michalopulos would be modified with the pixel-level feature vector object detection mechanism taught by Jeong resulting a in a system that can identify, virtualize, and control objects in an environment. Therefore, it would have been obvious to combine the teachings of the combination of Miller-Michalopulos and Jeong.
As per claim 3, Miller teaches, wherein the processor is further configured to:
receive a second image from the camera, wherein the second image shows a third physical device ([0032], the image sensors capture the images and/or live streams of the physical environment in real-time and provide the user with viewing 3D object-centric map representation of all physical objects as virtual objects or mixed reality objects on the display 108 of the AR device 102; and [0042], as the user moves throughout different spaces or zones or regions, artificial reality device 102 must provide synchronized, continuous, and updated feature maps with low latency in order to provide a high-quality, immersive, and enjoyable experience for users); and
in response to detecting that the third physical device has entered the boundary area, generate a second virtual device in the virtual environment, wherein the second virtual device is a virtual representation of the third physical device ([0004], each object is represented as a “virtual object” in the 3D map... the viewer is enabled to interact with each real object with the same experience as they experience during interacting within the real-world environment); and [0068], AR device 402 is configured to capture images of the scene of the living room 400 and detect each geometric dimension of the entire living room 400, including detecting the first features of each of the devices and objects...if the device and/or object are not matched with the predetermined features in the object and device library 126, then such devices and objects are determined to be new devices and objects. In such scenarios, the newly detected devices and objects are updated in the object and device library 126 for quick lookup next time when they are detected again. In some embodiments, the process of machine vision, artificial intelligence, deep learning, and machine learning helps in recognizing each device and each object accurately, automatically, and dynamically in real time. The identification of the devices and objects is used for generating map representation in the 3D map. In particular embodiments, the machine learning models are updated with each 3D map generated in each instance along with updating the instances or presence or existence of each device and object in a particular location in the physical environment).
Miller fails to specifically teach, determine a set of pixel location coordinates corresponding to a boundary area around the spatial computing device; determine a second pixel location coordinate associated with the third physical device in the second image; determine a second physical location coordinate of the third physical device in a global plane based at least in part upon the second pixel location coordinate associated with the third physical device; detect that the third physical device has entered the boundary area based at least in part upon the second physical location coordinate; and in response to detecting that the third physical device has entered the boundary area, generate a second virtual device in the virtual environment, wherein the second virtual device is a virtual representation of the third physical device.
However, Jeong teaches, determine a set of pixel location coordinates corresponding to a boundary area around the spatial computing device ([0012], the two-dimensional input image is obtained using a camera and includes two dimensional (2D) color data or grayscale data arranged in an array of pixels, each pixel corresponding to respective physical location within the scene; and [0045], Each image frame (also referred to as an “image” herein) can be organized as two-dimensional (2D) image data arranged as an X by Y array of picture element data structures (e.g., pixels) that are each indexed by a respective (x,y) coordinate pair, where each pixel stores image data that represents one or move values about light reflected from a corresponding point in an observed scene... pixel-level image data can take the form of a vector of elements that each represent a respective dimension or channel);
determine a second pixel location coordinate associated with the third physical device in the second image ([0045], Each image frame (also referred to as an “image” herein) can be organized as two-dimensional (2D) image data arranged as an X by Y array of picture element data structures (e.g., pixels) that are each indexed by a respective (x,y) coordinate pair, where each pixel stores image data that represents one or move values about light reflected from a corresponding point in an observed scene... pixel-level image data can take the form of a vector of elements that each represent a respective dimension or channel);
determine a second physical location coordinate of the third physical device in a global plane based at least in part upon the second pixel location coordinate associated with the third physical device ([0046], pixel locations in images captured for the different camera views can be mapped to physical locations within a scene that is captured in the images); and
detect that the third physical device has entered the boundary area based at least in part upon the second physical location coordinate ([0047], resulting point correspondence data can be applied to facilitate one or more of live detection, recognition and tracking operations in respect of the object 104).
The same motivation used in the rejection of claim 2 is applicable to the instant claim.
As per claim 4, Miller teaches, wherein establishing the communication path between the spatial computing device and the second physical device is in response to detecting that the second physical device has entered the boundary area ([0058], API platform 200 exposes the current eye vector from real-time eye tracking. The API platform 200 exposes the current object that the user is looking at or intending to interact with where such current object is detected through proximity and eye gaze).
As per claim 5, the combination of Miller-Michalopulos fails to specifically teach, wherein: the first set of features is represented by a first feature vector comprising a first set of numerical values; the second set of features is represented by a second feature vector comprising a second set of numerical values; and comparing each of the first set of features with the counterpart feature from among the second set of features comprises comparing each of the first set of numerical values with a counterpart number from among the second set of numerical values.
However, Jeong teaches, wherein: the first set of features is represented by a first feature vector comprising a first set of numerical values ([0058], Each pixel-level feature vector FV includes a respective set of generated features {f1, . . . , fn}, where n is the number of features (e.g., dimensions) per pixel; Examiner Note: feature vectors are arrays of numerical values);
the second set of features is represented by a second feature vector [comprising a second set of numerical values] ([0059], the reference point feature vector (RPFV) list 210 for a reference object 104R can, for example, include ... image coordinates for each of the selected reference points RP_1 to RP_Nrp (which can provide a reference point topography if required) and the respective reference feature vectors RF_1 to RFV_Nrp for each of the reference points RP_1 to RP_Nrp. Each reference feature vector RFV includes respective set of reference features {rf.sub.1, . . . , rf.sub.n} as generated by the trained feature generation model 124); and
comparing each of the first set of features with the counterpart feature from among the second set of features comprises comparing each of the first set of numerical values with a counterpart number from among the second set of numerical values ([0077], the feature vectors for pixels that map to common physical locations in a captured scene may be averaged across the view-specific feature maps 306 to provide average feature vectors that can then be compared to the reference feature vectors included in RPFV list 210, with a resulting POI list 210 being generated based on matching the average feature vectors to corresponding reference feature vectors. In other examples, the feature vectors for pixels that map to common physical locations from each of the multiple camera view images may each be independently compared to the reference feature vectors included in RPFV list 210 to identify points of interest in each of the multiple view images 302, with the POI lists for each set of multiple view images then used to generate a final POI list 310 for the set of multiple view images; and [0121], Every object can be detected in the input image/frame. By comparing the similarity scores of generated feature vectors, it is possible to find which one of the objects in the reference images is the most similar to the object of the input frame. In this way, the type of object in the input image/frame can be recognized).
The same motivation used in the rejection of claim 2 is applicable to the instant claim.
As per claim 9, this claim is similar to claim 2 and is rejected for the same reasons. The same motivation used in the rejection of claim 2 is applicable to the instant claim.
As per claim 10, this claim is similar to claim 3 and is rejected for the same reasons. The same motivation used in the rejection of claim 3 is applicable to the instant claim.
As per claim 11, this claim is similar to claim 4 and is rejected for the same reasons. The same motivation used in the rejection of claim 4 is applicable to the instant claim.
As per claim 12, this claim is similar to claim 5 and is rejected for the same reasons. The same motivation used in the rejection of claim 5 is applicable to the instant claim.
As per claim 16, this claim is similar to claim 2 and is rejected for the same reasons. The same motivation used in the rejection of claim 2 is applicable to the instant claim.
As per claim 17, this claim is similar to claim 3 and is rejected for the same reasons. The same motivation used in the rejection of claim 3 is applicable to the instant claim.
As per claim 18, this claim is similar to claim 4 and is rejected for the same reasons. The same motivation used in the rejection of claim 4 is applicable to the instant claim.
As per claim 19, this claim is similar to claim 5 and is rejected for the same reasons. The same motivation used in the rejection of claim 5 is applicable to the instant claim.
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure and is as follows:
O’Connell et al. (US 20230351747)- Discusses object detection using pixel-related data:
[0054], mage analysis engine 240 may identify the matching vehicle to the indicator generation engine 242. For example, the image analysis engine 240 may provide a location (e.g., pixel addresses, coordinates, etc.) in the image data corresponding to the matched vehicle; and [0151], the augmented reality module 1028 may capture image data using the camera 1044. The augmented reality module 1028 can determine whether a particular object (e.g., a hired vehicle, a particular passenger, a particular environment, etc.) is present in the image data, and overlay an indicator onto the image data to be displayed on the mobile device 1000
Ezrielev et al. (US 20250004463): Teaches virtualizing devices using a digital twin:
Abstract: simulating devices with a digital twin within the deployment. The digital twin may serve to replace operation of failed devices.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to MELISSA A HEADLY whose telephone number is (571)272-1972. The examiner can normally be reached Monday- Friday 9-5:30pm.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Bradley Teets can be reached at 571-272-3338. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/MELISSA A. HEADLY/
Examiner Art Unit 2197
/BRADLEY A TEETS/Supervisory Patent Examiner, Art Unit 2197