DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
This action is in response to the Amendment filed on 7/21/2026.
Claims 1-2, 5-12, 14-22, 23 are pending. Claims 1, 9, 10, 11, 17, 18 have been amended. Claim 3-4, 13 have been cancelled.
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claim(s) 1, 2, 5-12, 14-21, 23 is/are rejected under 35 U.S.C. 103 as being unpatentable over Pinheiro et al. (US 20180285686 A1, hereinafter Pinheiro), in view of Rabinovich et al. (US 20180053056 A1, hereinafter Rabinovich), further in view of Tveskov (US 20020140732 A1) and Szymczyk et al. (US 20120313969 A1, hereinafter Szymczyk).
Regarding Claim 10, Pinheiro teaches a system (Pinheiro, Fig. 11 Element System), comprising: at least one processor (Pinheiro, Fig. 11, Element 1102 Processor); and at least one memory (Pinheiro, Fig. 11, Element 1104 Memory) communicatively coupled to the at least one processor (Pinheiro, Fig. 11, Memory 1104 is communicatively coupled to the Processor 1102) and comprising computer-readable instructions that upon execution by the at least one processor cause the at least one processor to perform operations comprising (Pinheiro, Paragraph [0075], [0081], processor 1102 includes hardware for executing instructions, such as those making up a computer program. to execute instructions, processor 1102 may retrieve (or fetch) the instructions from an internal register, an internal cache, memory 1104, or storage 1106; a computer-readable non-transitory storage medium or media may include one or more semiconductor based or other integrated circuits (ICs)): extracting features from an image comprising an object by a computing device, wherein the features are extracted from the image using a first deep learning network model (Pinheiro, Paragraph [0047], “the system <read on computing device > may have a first, feature-extraction convolutional neural network (i.e., first convolutional neural network 510) that may take as inputs patches of images 410”; Fig. 9, Step 910, Paragraph [0065], the system processes a plurality of patches of an image, using a first deep-learning model, to detect a plurality of features associated with the first patch of the image”), wherein the object is associated with a geographic location and the first deep learning network model is pre-trained to extract features indicative of objects (Pinheiro, Paragraph [0047], “feature-extraction convolutional neural network (i.e., first convolutional neural network 510) that may take as inputs patches of images 410 and output features 520 of the patch/image (i.e., any number of features detected in the image)” “The feature-extraction layers may be pre-trained to perform classification on the image” “The feature-extraction model may be fine-tuned for object proposals during training of the system” [0065], “input the plurality of detected features associated with the respective patch of the image, and each object proposal includes a prediction as to a location of an object in the patch” [0035], “information of a concept may include …a location ( e.g., an address or a geographical location); determining one or more pre-stored files based on the geographic location of the object in the image (Pinheiro, Paragraph [0032], A privacy setting of a user determines how particular information associated with a user can be shared. Third-party-content-object stores may be used to store content objects received from third parties, such as a third-party system 170. Location stores may be used for storing location information received from client systems 130 associated with users; [0035], “information of a concept may include …a location ( e.g., an address or a geographical location; Paragraph [0043], “the system 400 is depicted as including a deep-learning model 420… the system may take in a plurality of patches of images 410 as inputs and output… for each patch input 410…identification of the location of the object)”), wherein each of the one or more pre-stored files comprises data indicative of a corresponding object (Pinheiro, Paragraph [0043], for each patch input 410, an object proposal 430 (i.e., a binary identification of the location of the object) and a score 440 (i.e., a scalar quantity predicting whether there is an object in the patch or not); recognizing the object based at least in part on the features extracted from the image (Pinheiro, Paragraph [0003], [0047], Convolutional neural networks may be used in large-scale object recognition tasks; The feature-extraction model may be fine-tuned for object proposals during training of the system). comparing the features extracted from the image with the data comprised in each of the one or more prestored files (Pinheiro, Abstract, Paragraph [0003], [0050], Convolutional neural networks may be used in large-scale object recognition tasks. the object-proposal branch may include a single 1×1 convolution layer followed by a classification layer (i.e., after the feature extraction layers of first convolutional neural network; “parameters while allowing each pixel classifier to leverage information from an entire feature map. The method enables generating the object proposal to provide information regarding a location of an object and to identify the object”); [[displaying an asset item in response to recognizing the object]]; and [[displaying a video stream and at least one selectable user interface element via an interface of the computing device in response to a request to trade the asset item for a different asset item, wherein the different asset item is displayed on a body part of a user associated with the computing device in the video stream, and wherein the at least one selectable user interface element is selectable to confirm or deny the requested trade]].
Pinheiro does not explicitly disclose but Rabinovich teaches displaying an asset item in response to recognizing the object (Rabinovich, Paragraph [0012], [0013], [0055], [0065], “FIG . 7 is a flowchart describing functioning of an example of an electromagnetic tracking system in the con text of an AR device” “FIG . 8 schematically illustrates examples of com ponents of an embodiment of an AR system” “a schematic illustrates coordination between the cloud computing assets (46) and local processing assets, which may, for example reside in head mounted componentry (58) coupled to the user's head“ “detection or calculation of head pose can facilitate the display system to render virtual objects”).
Rabinovich and Pinheiro are analogous since both of them are handling object with features in the augmented reality environment by using the deep learning network. Pinheiro provided a way of extracting the feature from objects and using the Convolutional Neural Network to train and learn the extracted representation of data. Rabinovich provided a way of extracting the feature from the object captured from the user connected camera including the asset saved in the augmented reality environment by using the Neural Network and to display the asset identified. Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention was made to incorporate saved asset handling taught by Rabinovich into modified invention of Pinheiro such that when processing the object in the augmented reality. System will be able to dynamically adjust the learning and training result by replying on the asset in the environment which captured from the user controlled camera and to display the asset in order to create more precise data representation when training the data in the augmented reality environment by using the Neural Network.
The combination of Pinheiro and Rabinovich does not explicitly disclose at least one selectable user interface element via an interface of the computing device in response to a request to trade the asset item for a different asset item, wherein the at least one selectable user interface element is selectable to confirm or deny the requested trade.
However, Tveskov teaches at least one selectable user interface element via an interface of the computing device in response to a request to trade the asset item for a different asset item, wherein the at least one selectable user interface element is selectable to confirm or deny the requested trade (Tveskov, Paragraph [0028], "Any user may initiate a trade sequence by selecting, for example, a trade button <read on selectable user interface element>. A selection allowing the user and trading partner to agree (e.g., agree buttons) <read on confirm the requested trade> to the trade may be included. Also, a selection for canceling the trade <read on deny the requested trade> may be included."; Paragraph [0055], "If the user wishes to make a trade for one or more trading icons 38 <read on asset item>, a "trade" icon 52 may be selected. If the trading pal responds with a "yes" icon, the user may be given the opportunity to accept <read on confirm the requested trade> or decline <read on deny the requested trade>"; Paragraph [0055], "FIG. 13 depicts an exemplary user interface for a request to trade <read on request to trade the asset item for a different asset item> along with a reply"; Paragraph [0055], "FIG. 14 depicts an exemplary user interface for accepting or canceling a trade <read on selectable user interface element selectable to confirm or deny the requested trade>").
Tveskov and Pinheiro are analogous since both of them relate to computer implemented systems for interacting with virtual objects through graphical user interfaces and networked environments. Pinheiro provides a way of extracting features from objects and recognizing objects using deep learning models. Tveskov provides a way of enabling users to exchange virtual objects through a graphical trading interface including selectable user interface elements for agreeing to or canceling a trade. Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention was made to incorporate the trade confirmation interface taught by Tveskov into the modified invention of Pinheiro such that users interacting with assets displayed in the augmented reality environment would be able to confirm or deny exchanges of the assets through selectable interface elements which improve user interaction and preventing unintended transfers of virtual assets.
The combination of Pinheiro, Rabinovich, Tveskov and Szymczyk does not explicitly disclose displaying a video stream in response to a request to trade the asset item for a different asset item, and does not explicitly disclose wherein the different asset item is displayed on a body part of a user associated with the computing device in the video stream.
However, Szymczyk teaches displaying a video stream (Szymczyk, Paragraph [0060], "the main display portion 204 may include a composite video feed <read on video stream> that incorporates a video feed of the user and one or more virtual-wearable items selected by the user via the item-search/selection module 118. The video feed of the user may be obtained by an imaging device associated with one of the client computing platforms 106"), wherein the different asset item is displayed on a body part of a user associated with the computing device in the video stream (Szymczyk, Paragraph [0062], "the motion-capture module 116 may be configured to recognize position and/or orientation of one more body parts of the user in the main display portion 204 in order to determine a position, size, and/or orientation for a given virtual-wearable item <read on different asset item> in the main display portion 204. Once the one or more body parts are recognized, the composite-imaging module 120 may position a virtual-wearable item at a predetermined offset and/or orientation relative to the recognized one or more body parts"; Paragraph [0063], "The motion-capture module 116 may track motion, position, and/or orientation of the user to overlay a virtual-wearable item on the user in the main display portion 204 such that the virtual-wearable item appears to be worn by the user while the user moves about and/or rotates in the main display portion 204").
Szymczyk and Tveskov are analogous since both relate to graphical user interfaces that allow a user to preview and interact with virtual items associated with the user before completing a transaction involving those items. Tveskov provides an interface that allows the user to select a trade of a virtual item and confirm or cancel that trade through selectable buttons, but its interface represents each item only as a static icon in a trading bin and does not visually preview the item on the user before the user commits. Szymczyk provides a virtual-outfitting interface whose main display portion contains a composite video feed incorporating a live video feed of the user with a selected virtual-wearable item overlaid on the user's recognized body parts, so that the user can see the item as it would appear worn on the user before deciding whether to act on it. Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to incorporate the composite-video-feed try-on interface taught by Szymczyk into the modified invention of Pinheiro such that, when Tveskov's trade request for a different asset item is received at the user's computing device, the interface additionally presents Szymczyk's composite video feed showing the different asset item overlaid on the recognized body parts of the user while the selectable agree/cancel controls of Tveskov remain available on the same interface. The motivation is to reduce the risk of buyer's remorse and unintended or unwanted trades, and to give the user a more informed and interactive basis on which to press the confirm or deny control, as Szymczyk itself explains at Paragraph [0004], [0003].
Regarding Claim 11, the combination of Pinheiro, Rabinovich, Tveskov and Szymczyk teaches the invention in claim 10.
The combination further teaches wherein the first deep learning network model is configured to be installed on computing device (Pinheiro, Paragraph [0004], [0023], A client system 130 may enable a network user at client system 130 to access network 110. a system may use one or more deep-learning models to generate a number of object proposals (i.e., masks) for an image).
Regarding Claim 12, the combination of Pinheiro, Rabinovich, Tveskov and Szymczyk teaches the invention in Claim 10.
The combination further teaches determining the geographic location based on
information indicating a position of a camera of a client computing device, wherein the
information indicating the position of the camera comprises GPS (Global Position
System) information (Pinheiro, Paragraph [0032], "Location stores may be used for
storing location information <read on GPS (Global Position System)> received
from client systems 130 associated with users."; Paragraph [0035], "a location (e.g., an address or a geographical location); it is noted location information received from client system associated with users device like mobile telephone through communication network which carry satellite communication which can carry GPS signal).)
But Pinheiro does not explicitly disclose that the system determines the geographic location based on the camera position information itself.
However, Rabinovich teaches determining the geographic location based on
information indicating a position of a camera of a client computing device (Rabinovich,
Paragraph [0012], [0013], "FIG. 7 is a flowchart describing functioning of an
example of an electromagnetic tracking system in the context of an AR device.";
"FIG. 8 schematically illustrates examples of components of an embodiment of an
AR system." Paragraph [0124], [0068], [0161], the pose (e.g., vector and/or origin position information relative to the world) of the cameras that capture those images or points may be determined. The result may be that the 3-D points and also the calculated trajectory (e.g., location, path of the capturing cameras) may be adjusted by a small amount; after a user powers up his or her wearable computing system (160), a head mounted component assembly may capture a combination of IMU and camera dataIt is noted that the AR system uses tracking data to determine camera pose and position which is determining the geographic location based on camera position.)
Pinheiro, Rabinovich, Tveskov and Szymczyk are analogous since both are dealing with systems that
associate captured image content with geographic context and use camera-based data
in an augmented-reality environment. Pinheiro provided a way of storing and using GPS-based location information for objects and users in its social-networking and image-analysis system. Rabinovich provided a way of determining the camera's pose and position from tracking systems to compute the device's geographic location for AR display alignment. Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to incorporate the camera-position determination taught by Rabinovich into the modified invention of Pinheiro such that the system in Pinheiro determines its geographic location based on camera position information, thereby improving accuracy and contextual alignment of displayed objects.
Regarding Claim 14, the combination of Pinheiro, Rabinovich, Tveskov and Szymczyk teaches the invention in claim 10.
The combination further teaches wherein each of the one or more pre-stored files comprises features extracted from one or more images comprising the corresponding object (Pinheiro, Paragraph [0003], [0047], A system may use object-identification algorithms to identify, for an object proposal, what the corresponding object is. The feature-extraction layers may be pre-trained to perform classification on the image. The feature-extraction model may be fine-tuned for object proposals during training of the system), and the features are extracted from the one or more image using a second deep learning network model (Pinheiro, Paragraph [0065], a plurality of features associated with the first patch of the image. Each patch includes one or more pixels of the image. At step 920, the system generates, using a second deep-learning model. The second deep-learning model takes as input the plurality of detected features associated with the respective patch of the image, and each object proposal includes a prediction as to a location of an object in the patch).
Regarding Claim 15, the combination of Pinheiro, Rabinovich, Tveskov and Szymczyk teaches the invention in claim 10.
The combination further teaches wherein a plurality of sets of pre-stored files are associated with a plurality of geographic locations (Pinheiro, Paragraph [0035], [0073], appropriate, computer system 1100 may include one or more computer systems 1100; be unitary or distributed; span multiple locations; One or more computer systems 1100 may perform at different times or at different locations one or more steps…a location ( e.g., an address or a geographical location)).
Regarding Claim 16, the combination of Pinheiro, Rabinovich, Tveskov and Szymczyk teaches the invention in claim 10.
The combination further teaches the operations further comprising: storing data indicative of the asset item in response to user input (Rabinovich, Paragraph [0065], [0066], These computing assets local to the user may be operatively coupled to each other as well, via wired and/or wireless; a map of the world may be continually updated at a storage location which may partially reside on the user's AR system).
As explained in rejection of claim 10, the obviousness for combining of asset of Rabinovich into Pinheiro is provided above.
Regarding Claim 17, the combination of Pinheiro, Rabinovich, Tveskov and Szymczyk teaches the invention in claim 10.
The combination further teaches the operations further comprising: determining the body part of the user in the video stream (Szymczyk, Paragraph [0062], "the motion-capture module 116 may be configured to recognize <read on determining> position and/or orientation of one more body parts of the user in the main display portion 204 <read on in the video stream — the main display portion carries the composite video feed per Paragraph [0060]> in order to determine a position, size, and/or orientation for a given virtual-wearable item in the main display portion 204. Once the one or more body parts are recognized, the composite-imaging module 120 may position a virtual-wearable item at a predetermined offset and/or orientation relative to the recognized one or more body parts"; Paragraph [0063], "The motion-capture module 116 may track motion, position, and/or orientation of the user to overlay a virtual-wearable item on the user in the main display portion 204 such that the virtual-wearable item appears to be worn by the user while the user moves about and/or rotates in the main display portion 204"); and displaying an effect of the asset item being tried on the body part of the user in the video stream (Szymczyk, Paragraph [0018], "The user may select one or more virtual-wearable items from to queue in order to virtually "try on" real-wearable items corresponding to selected virtual-wearable items"; Paragraph [0060], "The main display portion 204 may include one or more images and/or video of the user virtually trying on one or more real-wearable items that correspond to one or more selected virtual-wearable items. In such images and/or video, the one or more selected virtual-wearable items may be visually overlaid on the user in a position in which the user would normally wear corresponding real-wearable items"; Paragraph [0063], "The motion-capture module 116 may track motion, position, and/or orientation of the user to overlay a virtual-wearable item <read on asset item> on the user in the main display portion 204 such that the virtual-wearable item appears to be worn <read on the effect of the asset item being tried on> by the user while the user moves about and/or rotates in the main display portion 204 <read on in the video stream>").
As explained in rejection of claim 1, the obviousness for combining of Szymczyk into Pinheiro is provided above.
Regarding Claim 1, it recites limitations similar in scope to the limitations of Claim 10 but as a method and the combination of Pinheiro, Rabinovich, Tveskov and Szymczyk teaches all the limitations as of Claim 10. Therefore is rejected under the same rationale.
Regarding Claim 2, it recites limitations similar in scope to the limitations of Claim 11 and therefore is rejected under the same rationale.
Regarding Claim 5, it recites limitations similar in scope to the limitations of Claim 14 and therefore is rejected under the same rationale.
Regarding Claim 6, it recites limitations similar in scope to the limitations of Claim 15 and therefore is rejected under the same rationale.
Regarding Claim 7, the combination of Pinheiro, Rabinovich, Tveskov and Szymczyk teaches the invention in claim 1.
The combination further teaches wherein the object comprises a unique immobile object (Rabinovich, Paragraph [0163], [0192], “The middle layers (266) may be configured to start learning parts, for example—object parts, face features, and the like;” “During the vision pose calculation process , there is an assumption that features being viewed by the outward facing cameras are static features ( e . g., not moving from frame to frame relative to the global coordinate system)”; it is noted since the object is not moving it is immobile).
Regarding Claim 8, it recites limitations similar in scope to the limitations of Claim 16 and therefore is rejected under the same rationale.
Regarding Claim 9, it recites limitations similar in scope to the limitations of Claim 17 and therefore is rejected under the same rationale.
Regarding Claim 18, it recites limitations similar in scope to the limitations of claim 10 and the combination of Pinheiro, Rabinovich, Tveskov and Szymczyk teaches all the limitations as of Claim 10. And Pinheiro discloses these features can be implemented on a computer-readable storage medium (Pinheiro, Paragraph [0075], [0081], “a computer-readable non-transitory storage medium or media may include one or more semi-conductor based or other integrated circuits” “processor 1102 includes hardware for executing instructions, such as those making up a computer program”)
Regarding Claim 19, the combination of Pinheiro, Rabinovich, Tveskov and Szymczyk teaches the invention in Claim 18.
The combination further teaches wherein the first deep learning network model is configured to be installed on the computing device (Pinheiro, Paragraph [0004], [0023], A client system 130 may enable a network user at client system 130 to access network 110. a system may use one or more deep-learning models to generate a number of object proposals (i.e., masks) for an image), wherein each of the one or more pre-stored files comprises features extracted from one or more images comprising the corresponding object(Pinheiro, Paragraph [0003], [0047], A system may use object-identification algorithms to identify, for an object proposal, what the corresponding object is. The feature-extraction layers may be pre-trained to perform classification on the image. The feature-extraction model may be fine-tuned for object proposals during training of the system),
and the features are extracted from the one or more image using a second deep learning network model (Pinheiro, Paragraph [0065], a plurality of features associated with the first patch of the image. Each patch includes one or more pixels of the image. At step 920, the system generates, using a second deep-learning model. The second deep-learning model takes as input the plurality of detected features associated with the respective patch of the image, and each object proposal includes a prediction as to a location of an object in the patch),
and the features are extracted from the one or more image using a second deep learning network model (Pinheiro, Paragraph [0065], a plurality of features associated with the first patch of the image. Each patch includes one or more pixels of the image. At step 920, the system generates, using a second deep-learning model. The second deep-learning model takes as input the plurality of detected features associated with the respective patch of the image, and each object proposal includes a prediction as to a location of an object in the patch);
Regarding Claim 20, it recites limitations similar in scope to the limitations of Claim 17 and therefore is rejected under the same rationale.
Regarding Claim 21, the combination of Pinheiro, Rabinovich, Tveskov and Szymczyk teaches the invention in Claim 1.
The combination further teaches determining the geographic location based on information indicating a position of a camera of a client computing device (Rabinovich,
Paragraph [0012], [0013], "FIG. 7 is a flowchart describing functioning of an
example of an electromagnetic tracking system in the context of an AR device.";
"FIG. 8 schematically illustrates examples of components of an embodiment of an
AR system." Paragraph [0124], [0068], [0161], the pose (e.g., vector and/or origin position information relative to the world) of the cameras that capture those images or points may be determined. The result may be that the 3-D points and also the calculated trajectory (e.g., location, path of the capturing cameras) may be adjusted by a small amount; after a user powers up his or her wearable computing system (160), a head mounted component assembly may capture a combination of IMU and camera dataIt is noted that the AR system uses tracking data to determine camera pose and position which is determining the geographic location based on camera position.), wherein the image is captured by the camera associated with a user (Rabinovich, Paragraph [0066], As more and more AR users continually capture information about their real environment (e.g., through cameras, sensors, IMUs, etc.).
As explained in rejection of claim 1, the obviousness for combining of asset of Rabinovich into Pinheiro is provided above.
Regarding Claim 23, the combination of Pinheiro, Rabinovich, Tveskov and Szymczyk teaches the invention in claim 21.
The combination further teaches wherein the information indicating the position of the camera comprises GPS (Global Position System) information (Pinheiro, Paragraph
[0022], “Links 150 may connect client system 130, social networking system 160, and third-party system 170 to communication network 110 or to each other… In particular embodiments, one or more links 150 include one or more… a satellite communications technology based network” [0032], "Location stores may be used for storing location information received from client systems 130 associated with users."; Paragraph [0035], "a location (e.g., an address or a geographical location); it is noted location information received from client system associated with users device like mobile telephone through communication network which carry satellite communication which can carry GPS signal).
Claim(s) 22 is/are rejected under 35 U.S.C. 103 as being unpatentable over Pinheiro et al. (US 20180285686 A1, hereinafter Pinheiro), in view of Rabinovich et al. (US 20180053056 A1, hereinafter Rabinovich), further in view of Tveskov (US 20020140732 A1) and Szymczyk et al. (US 20120313969 A1, hereinafter Szymczyk) as applied to Claim 1 above and further in view of Mallett et al. (US 20210398095 A1, hereinafter Mallett)
Regarding Claim 22, the combination of Pinheiro, Rabinovich, Tveskov and Szymczyk teaches the invention in Claim 1.
The combination does not explicitly disclose but Mallott teaches wherein the asset item is a token representing an association between the user and the recognized object (Mallott, Paragraph [0002], Such digital environments are provided by content providers. Example digital environments include social networks [0010], the branded digital item is a recognizable graphical object; generates a branded digital item blockchain that includes a non-fungible token associated with the branded digital item; receives information corresponding to a purchase of the branded digital item by a user; and updates the non-fungible token to memorialize purchase of the branded digital item by the authorized user).
Mallott and Pinheiro are analogous since both of them are handling object with features in the social networking environment. Pinheiro provided a way of extracting the feature from objects and using the Convolutional Neural Network to train and learn the extracted representation of data. Mallott provided a way of extracting the feature from the object captured by using token to relating user and the object in the social networking environment. Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention was made to incorporate token relationship taught by Mallott into modified invention of Pinheiro such that when processing the object in the augmented reality. System will be able to vdynamically adjust the learning and training result by replying on the asset in the environment which linking between user and the object using token which increase the accuracy and to create more precise data representation when training the data in the augmented reality environment by using the Neural Network.
Response to Arguments
Applicant’s arguments with respect to claim 1, 10, 18, filed on 7/21/2026, with respect to rejection under 35 USC § 103 in regard to prior art does not teaches the limitation(s) “wherein the different asset item is displayed on a body part of a user associated with the computing device in the video stream" have been considered but is in moot since it has now been taught by the combination of Pinheiro, Rabinovich, Tveskov and Szymczyk.
Applicant further asserts that the prior art does not teach the limitation "recognizing the object based at least in part on comparing the features extracted from the image with data comprised in the one or more pre-stored files," arguing that Pinheiro's trained convolutional-neural-network weights and biases "are not the same as pre-stored object files" and "do not store object data or correspond to specific objects in the same way that a file containing data about an object would."
In response to the argument, as described in the rejection of claim 1 above, this limitation is taught by Pinheiro under the broadest reasonable interpretation of the claim language. The claim language "one or more pre-stored files comprises data indicative of a corresponding object" is broad on its face — the claim does not require any particular file format, does not require one file per object, does not require raw image templates or per-object reference images, and does not exclude data that has been encoded, distilled, or otherwise learned from training. Any stored data structure that indicates the presence or identity of a corresponding object satisfies the claim. Applicant's argument mischaracterizes what Pinheiro actually stores. Pinheiro is not limited to the trained weights of a classification layer. Pinheiro expressly teaches distinct, per-object pre-stored data structures that are indicative of corresponding objects: Pinheiro, Paragraph [0032], teaches "Third-party-content-object stores may be used to store content objects received from third parties, such as a third-party system 170. Location stores may be used for storing location information received from client systems 130 associated with users." These are pre-stored files, keyed to content objects and locations, that contain data indicative of corresponding objects. This teaching sits squarely within the "one or more pre-stored files" language of the claim. Pinheiro, Paragraph [0035], teaches "information of a concept may include …a location (e.g., an address or a geographical location)." Each such "concept" is per-object data indicative of the corresponding object, stored in a manner that can be looked up by geographic location — which is precisely the manner in which the claim recites the pre-stored files being determined ("determining one or more pre-stored files based on the geographic location of the object in the image"). Pinheiro, Paragraph [0043], teaches "for each patch input 410, an object proposal 430 (i.e., a binary identification of the location of the object) and a score 440 (i.e., a scalar quantity predicting whether there is an object in the patch or not)." The system compares extracted features against this per-object reference data to produce the object identification. Pinheiro, Paragraph [0050], teaches "The method enables generating the object proposal to provide information regarding a location of an object and to identify the object." The identification of the object is performed by comparing extracted features against the stored object-indicative data described in the paragraphs above; without such stored data indicative of objects, no identification could be produced. Applicant's argument treats the trained CNN weights in isolation and argues that they are only numerical parameters. But Pinheiro's disclosure is not so limited. The content-object stores, location stores, and per-concept information described in Paragraphs [0032], [0035], and [0043] are separate, per-object stored data structures that contain data indicative of corresponding objects. The claim recites data "indicative of a corresponding object" — not "identical to" or "consisting of raw image data of" the corresponding object — and Pinheiro's stored data structures satisfy that requirement. The recognition operation of Pinheiro compares the extracted features against these stored, per-object references to produce the identification, precisely as the claim requires. Furthermore, Applicant's characterization that the CNN's learned parameters "do not correspond to specific objects" is not supported by Pinheiro's own disclosure. Pinheiro, Paragraph [0047], teaches "The feature-extraction layers may be pre-trained to perform classification on the image" and "The feature-extraction model may be fine-tuned for object proposals during training of the system." The learned parameters are trained to classify and propose specific objects — that is by definition data that is "indicative of a corresponding object," even if it has been encoded through a learning process rather than stored as raw templates. The claim does not preclude the pre-stored files from being trained or learned representations. Hence the combination of prior arts (Pinheiro in view of Rabinovich, Tveskov, and Szymczyk) fully teaches the limitations of independent claims 1, 10, and 18, including the "recognizing the object based at least in part on comparing the features extracted from the image with the data comprised in each of the one or more pre-stored files" limitation. Therefore Applicant's remark cannot be considered persuasive.
In regard to Claims 2, 5-9, 11-12, 14-17, 19-23, they directly/indirectly depends on independent Claim 1, 10, 18 respectively. Applicant does not argue anything other than the independent claim 1, 10, 18. The limitations in those claims in conjunction with combination previously established as explained.
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure.
US 20190171970 A1 Trading goods based on image processing for interest, emotion and affinity detection
Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a).
A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to YUJANG TSWEI whose telephone number is (571)272-6669. The examiner can normally be reached 8:30am-5:30pm EST.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Kent Chang can be reached at (571)272-7667. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/YuJang Tswei/Primary Examiner, Art Unit 2614