DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Claim Rejections - 35 USC § 112
The following is a quotation of 35 U.S.C. 112(b):
(b) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention.
The following is a quotation of 35 U.S.C. 112 (pre-AIA ), second paragraph:
The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the applicant regards as his invention.
Claims 1-20 are rejected under 35 U.S.C. 112(b) or 35 U.S.C. 112 (pre-AIA ), second paragraph, as being indefinite for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor (or for applications subject to pre-AIA 35 U.S.C. 112, the applicant), regards as the invention.
Claim 1 recites “image data” in lines 2 and 6. It is unclear to the Examiner whether the limitation in line 2 is the same or different from the limitation in line 6.
Claim 1 recites “a plurality of tags”, “tags”, and “a respective tag” in lines 14, 19, 21-22, respectively. It is unclear to the Examiner whether the limitations are the same or different from each other.
Claim 5 recites “the second frame” in line 4. There is insufficient antecedent basis for this limitation in the claim.
Claims 2-4, 6-10 depend on at least claim 1. Therefore, the claims 2-4, 6-10 are rejected for at least the same reason as claim 1.
Claims 11, 17 recite similar limitations as claim 1. Therefore, the claims 11 and 17 require similar corrections as claim 1.
Claim 16 recites “the second frame” in line 3. There is insufficient antecedent basis for this limitation in the claim.
Claims 12-15 depend on at least claim 11. Therefore, the claims 12-15 are rejected for at least the same reason as claim 11.
Claims 18-20 depend on at least claim 17. Therefore, the claims 18-20 are rejected for at least the same reason as claim 17.
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows:
1. Determining the scope and contents of the prior art.
2. Ascertaining the differences between the prior art and the claims at issue.
3. Resolving the level of ordinary skill in the pertinent art.
4. Considering objective evidence present in the application indicating obviousness or nonobviousness.
Claim(s) 1, 5, 9, 10-12, 14, 15, 17, 18 is/are rejected under 35 U.S.C. 103 as being unpatentable over Borhan et al. (U.S. Patent Application 20130218721) in view of Barnett et al. (U.S. Patent Application 20180190032).
In regards to claim 1, Borhan teaches a computer-implemented method [e.g. an augmented retail shopping processor-implemented method, 0360], the method comprising:
obtaining, by a computing system [Fig. 2C-2D; e.g. TVC system, 0086] comprising one or more processors [e.g. processor, 0324], image data [Fig. 5C, 20D; e.g. capturing live video of the virtual reality scene, 0190, also see 0105], wherein the image data comprises a plurality of image frames [e.g. the TVC may obtain two consecutive video frame grabs 2071 (e.g., every 100 ms, etc.), 0187];
determining, by the computing system, a first image frame and a second image frame are associated with a scene [e.g. The TVC may compare the two consecutive video frames 2075 (e.g., via histogram comparison, etc.), and determine the difference region of the two frames. The cameras 532a may be positioned in a grid such that the visual scope 532b of the cameras overlap, allowing TVC to stitch together images to create a panoramic view of the store aisle, 0187, also see 0117];
generating, by the computing system, image data comprising the first image frame and the second image frame of the plurality of image frames [e.g. the virtual store may be comprised of stitched-together composite photographs having detailed GPS coordinates related to each individual photograph and having detailed accelerometer gyroscopic, positional/directional information, all of which may be used to allow TVC to stitch together a virtual and continuous composite view of the store, 0116];
processing, by the computing system, the image data with a pattern recognition model to generate object data descriptive of a plurality of objects determined to be in the scene [e.g. The TVC may then perform OCR and/or pattern recognition on the obtained image (e.g., around the fingertip position) 2008 to determine a type of the object in the image 2010, 0177];
obtaining, by the computing system and by performing a plurality of searches based on the object data, object-specific information for at least a subset of the plurality of objects, wherein the object-specific information comprises one or more details for each of the at least the subset of the plurality of objects [e.g. The TVC may query the obtained product inventory and stock keeping data based on the product identifier and the product category for each product item, and determine an in-store stock keeping location for each product item based on the query. A social layer 715d to obtain social rating/review information, such as Amazon ratings, Facebook comments, Tweets, related products, friends ratings, top reviews, and/or the like, 0113, 0125];
generating, by the computing system, a plurality of tags based on the object-specific information [e.g. the TVC may retrieve a virtual label template 2061 based on the information type and populate relevant information into the label template 2062. The TVC may provide a tag label overlaid on top of the item showing product information 1607, 0191, also see 0160]; and
providing, by the computing system, a plurality of user-interface elements associated with the plurality of tags rendered within one or more images of the plurality of image frames via an augmented-reality experience, wherein a subset of the plurality of user-interface elements comprise tags overlaid over respective objects [Fig. 5C; e.g. The virtual label may be positioned close to the object, and inject the generated virtual label overlaying the live video at the position 2065. Virtual overlay labels such as Apple Jam, Whipped Cream, Baking Soda are overlaid over the product items, 0191, also see 0115], and wherein one or more of the user-interface elements of the plurality of user-interface elements comprise one or more off-screen indicators that indicate a respective object of the plurality of objects is not currently displayed and has a respective tag [Fig. 5C; e.g. The virtual overlay labels provide directions for the consumer to locate other product items that are not located within the captured reality scene 516, 0115].
Borhan does not explicitly teach
processing, by the computing system, the image data with a machine-learned recognition model to generate object data descriptive of a plurality of objects determined to be in the scene (emphasis added).
However, Barnett teaches
processing, by the computing system, the image data [e.g. image data of the user’s camera view, 0038] with a machine-learned recognition model [e.g. one or more machine learning models can be trained to identify objects depicted in a user's camera view, 0044] to generate object data [e.g. location information, 0047] descriptive of a plurality of objects determined to be in the scene [e.g. one or more objects depicted in a user's camera view, 0047].
Therefore, it would have been obvious to one of ordinary skill in the art to have modified Borhan’s method with the features of
processing, by the computing system, the image data with a machine-learned recognition model to generate object data descriptive of a plurality of objects determined to be in the scene
in the same conventional manner as taught by Barnett because Barnett provides an efficient method of quickly identifying objects automatically [0047].
In regards to claim 5, Borhan teaches the method of claim 1, wherein generating, by the computing system, image data comprising the first image frame and the second image frame of the plurality of image frames [see rejection of claim 1 above] comprises:
stitching together the first image frame and the second frame of the plurality of image frames to generate a stitched image [e.g. the cameras 532a may be positioned in a grid such that the visual scope 532b of the cameras overlap, allowing TVC to stitch together images to create a panoramic view of the store aisle, 0117].
In regards to claim 9, Borhan teaches the method of claim 5, wherein the stitched image is generated with a stitching model [Fig. 44; e.g. image processing component 4443, 0345, also see 0116-0117].
In regards to claim 10, Borhan teaches the method of claim 9, wherein the stitching model:
processes the plurality of image frames to determine two or more image frames are descriptive of a same scene [e.g. The cameras 532a may be positioned in a grid such that the visual scope 532b of the cameras overlap, allowing TVC to stitch together images to create a panoramic view of the store aisle. The virtual store may be comprised of stitched-together composite photographs having detailed GPS coordinates related to each individual photograph and having detailed accelerometer gyroscopic, positional/directional information, all of which may be used to allow TVC to stitch together a virtual and continuous composite view of the store, 0116-0117]; and
in response to determining the two or more image frames are descriptive of the same scene, generates scene data descriptive of the image frames being stitched together [e.g. the visual scope 532b of the cameras overlap, allowing TVC to stitch together images to create a panoramic view of the store aisle, 0117].
In regards to claim 11, the claim recites similar limitations as claim 1, but in the form of a computing system, the system comprising: one or more processors; and one or more non-transitory computer-readable media that collectively store instructions that, when executed by the one or more processors, cause the computing system to perform the method of claim 1. Furthermore, Borhan teaches a computing system [Fig. 2C-2D; e.g. TVC system, 0086], the system comprising: one or more processors [e.g. processor, 0324]; and one or more non-transitory computer-readable media [e.g. memory, 0324] that collectively store instructions [e.g. collection of program components, 0325] that, when executed by the one or more processors, cause the computing system to perform the method of claim 1. Therefore, the same rationale as claim 1 is applied.
In regards to claim 12, Borhan teaches the system of claim 11, wherein determining the first image frame and the second image frame are associated with the scene comprises: determining the first image frame and the second image frame were captured at a particular location [e.g. the virtual store may be comprised of stitched-together composite photographs having detailed GPS coordinates related to each individual photograph and having detailed accelerometer gyroscopic, positional/directional information, all of which may be used to allow TVC to stitch together a virtual and continuous composite view of the store, 0116].
In regards to claim 14, Borhan teaches the system of claim 12, wherein the particular location is determined based on one or more location sensors on a user computing device that captured the plurality of image frames [e.g. The mobile device 813 may capture the consumer's GPS coordinates 826. The detailed GPS coordinates related to each individual photograph and having detailed accelerometer gyroscopic, positional/directional information, 0130, also see 0116].
In regards to claim 15, Borhan teaches the system of claim 11, wherein the plurality of tags comprise a plurality of candidate queries generated based on the object-specific information for at least the subset of the plurality of objects [e.g. the TVC may provide identified item information 1631 in a virtual label, and alternative item recognition information 1632, 1633, 1634 and/or provide an option to the consumer to see more similar products 1635, 0167].
In regards to claim 17, the claim recites similar limitations as claim 1, but in the form of a one or more non-transitory computer-readable media that collectively store instructions that, when executed by one or more computing devices, cause the one or more computing devices to perform the method of claim 1. Furthermore, Borhan teaches a one or more non-transitory computer-readable media [e.g. memory, 0324] that collectively store instructions [e.g. collection of program components, 0325] that, when executed by one or more computing devices [Fig. 2C-2D; e.g. TVC system, 0086], cause the one or more computing devices to perform the method of claim 1. Therefore, the same rationale as claim 1 is applied.
In regards to claim 18, Borhan teaches the one or more non-transitory computer-readable media of claim 17, wherein the respective tag is descriptive of a distinguishing feature of the respective object [e.g. a user may place two payment cards in the scene so that the TVC may capture the cards. For example, the TVC may capture the type of the card, e.g., Visa 1608a and MasterCard 1608b, and provide labels to show rebate/rewards policy associated with each card for such a transaction 1609a-b. As such, the user may select to pay with a card to gain the provided rebate/rewards, 0161].
Claim(s) 7 is/are rejected under 35 U.S.C. 103 as being unpatentable over Borhan et al. (U.S. Patent Application 20130218721) in view of Barnett et al. (U.S. Patent Application 20180190032) as applied to claims 5 above, and further in view of Shi et al. (U.S. Patent Application 20150248591).
In regards to claim 7, Borhan as modified by Barnett does not explicitly teach the method of claim 5, further comprising: providing, by the computing system, the stitched image for display with the plurality of tags.
However, Shi teaches the method of claim 5, further comprising: providing, by the computing system, the stitched image for display with the plurality of tags [e.g. The image recognition system 204 outputs the merged recognition results and a single stitched image. Each recognition result may include a list of descriptions of the recognized objects from an input image. The description for each recognized object may include, for example, an object label, an object ID (e.g., a stock keeping unit (SKU)), and coordinates of a bounding box indicating where the object is located in the input image. The description for each recognized object may also include other information including a confidence of the recognition module 301 in identifying the object, 0045].
Therefore, it would have been obvious to one of ordinary skill in the art to have modified the combination of Borhan’s method and the teachings of Barnett with the features of providing, by the computing system, the stitched image for display with the plurality of tags in the same conventional manner as taught by Shi because Shi provides a method for capturing multiple images of shelves and recognizing as many products and the locations of those products as possible by not to double counting products that appear in multiple images [0007-0008].
Claim(s) 16 is/are rejected under 35 U.S.C. 103 as being unpatentable over Borhan et al. (U.S. Patent Application 20130218721) in view of Barnett et al. (U.S. Patent Application 20180190032) as applied to claims 11 above, and further in view of Fischer et al. (FlowNet: Learning Optical Flow with Convolutional Networks).
In regards to claim 16, Borhan as modified by Barnett does not explicitly teach the system of claim 11, wherein generating the image data comprising the first image frame and the second image frame of the plurality of image frames comprises: concatenating the first image frame and the second frame of the plurality of image frames.
However, Fischer teaches the system of claim 11, wherein generating the image data comprising the first image frame and the second image frame of the plurality of image frames [see rejection of claim 11] comprises: concatenating the first image frame and the second frame of the plurality of image frames [e.g. stack both input images together, see section titled “Network Architectures” in page 3].
Therefore, it would have been obvious to one of ordinary skill in the art to have modified the combination of Borhan’s method and the teachings of Barnett with the features of concatenating the first image frame and the second frame of the plurality of image frames in the same conventional manner as taught by Fischer because concatenating two images is well known and commonly used in the art of image processing systems.
Claim(s) 19 is/are rejected under 35 U.S.C. 103 as being unpatentable over Borhan et al. (U.S. Patent Application 20130218721) in view of Barnett et al. (U.S. Patent Application 20180190032) as applied to claims 17 above, and further in view of Bathiche et al. (U.S. Patent 8,687,021).
In regards to claim 19, Borhan as modified by Barnett does not explicitly teach the one or more non-transitory computer-readable media of claim 17, wherein the plurality of tags are ranked and selected based on a scene context.
However, Bathiche teaches the one or more non-transitory computer-readable media of claim 17, wherein the plurality of tags are ranked and selected based on a scene context [e.g. The innovation can also filter, rank, modify or ignore virtual-world information based upon a particular real-world class, user identity or context, see Abstract].
Therefore, it would have been obvious to one of ordinary skill in the art to have modified the combination of Borhan’s method and the teachings of Barnett with the features of wherein the plurality of tags are ranked and selected based on a scene context in the same conventional manner as taught by Bathiche because Bathiche provides a method to enhance the visual experience by automatically and dynamically overlaying or interspersing virtual capabilities (e.g., data) with real world situations [c.1 L.53-65].
Claim(s) 20 is/are rejected under 35 U.S.C. 103 as being unpatentable over Borhan et al. (U.S. Patent Application 20130218721) in view of Barnett et al. (U.S. Patent Application 20180190032) as applied to claims 17 above, and further in view of Tiutiunnik et al. (U.S. Patent 11,430,211).
In regards to claim 20, Borhan as modified by Barnett does not explicitly teach the one or more non-transitory computer-readable media of claim 17, wherein the plurality of tags are selected based on tag popularity among a plurality of other users.
However, Tiutiunnik teaches the one or more non-transitory computer-readable media of claim 17, wherein the plurality of tags are selected based on tag popularity among a plurality of other users [e.g. the display priority can be computed based on the following parameters: tag rating, and also tag associations with other users, or a combination of the above parameters, c.10 L.32-54, c.10 L.66-c.11 L.11].
Therefore, it would have been obvious to one of ordinary skill in the art to have modified the combination of Borhan’s method and the teachings of Barnett with the features of wherein the plurality of tags are selected based on tag popularity among a plurality of other users in the same conventional manner as taught by Tiutiunnik because Tiutiunnik provides a method to avoid the inaccessibility of social media information relevant to a particular real-world object, connect users based on what objects they are interested in, improve user experience while communicating in social media, as well as increase communication effectiveness (for those purposes, a number of visually distinctive real-world phenomena can be utilized in the same way as objects) [c.1 L.37-45].
Allowable Subject Matter
Claims 2-4, 6, 8, 13 would be allowable if rewritten to overcome the rejection(s) under 35 U.S.C. 112(b) or 35 U.S.C. 112 (pre-AIA ), 2nd paragraph, set forth in this Office action and to include all of the limitations of the base claim and any intervening claims.
In regards to claim 2, the prior art of record fails to teach or suggest the method of claim 1, wherein generating the plurality of tags comprises:
determining, by the computing system, a plurality of differentiating attributes associated with differentiators between the at least the subset of the plurality of objects; and
generating, by the computing system, the plurality of tags based on the plurality of differentiating attributes.
In regards to claim 3, the prior art of record fails to teach or suggest the method of claim 1, further comprising:
processing, by the computing system, object-specific information for the plurality of objects and the image data to determine a plurality of filters; and
providing, by the computing system, one or more selectable user-interface elements overlaid over the image data, wherein the one or more selectable user-interface elements are descriptive of one or more particular filters of the plurality of filters.
In regards to claim 4, the claim depends on at least claim 3. Therefore, the claim is allowable for at least the same reason as claim 3 if rewritten to overcome the rejection(s) under 35 U.S.C. 112(b) or 35 U.S.C. 112 (pre-AIA ), 2nd paragraph, set forth in this Office action and to include all of the limitations of the base claim and any intervening claims.
In regards to claim 6, the prior art of record fails to teach or suggest the method of claim 5, wherein generating, by the computing system, image data comprising the first image frame and the second image frame of the plurality of image frames further comprises: cropping the stitched image to remove data that is irrelevant to a semantic understanding of the scene.
In regards to claim 8, the prior art of record fails to teach or suggest the method of claim 5, wherein processing, by the computing system, the image data with the machine-learned recognition model to generate the object data descriptive of the plurality of objects determined to be in the scene comprises:
processing the stitched image with the machine-learned recognition model to generate the object data descriptive of the plurality of objects determined to be in the scene, wherein the stitched image is processed without providing the stitched image for display.
In regards to claim 13, the prior art of record fails to teach or suggest the system of claim 12, wherein the particular location is determined based on the time between image frames being below a threshold time.
Conclusion
Any inquiry concerning this communication or earlier communications from the examiner should be directed to ANDREW SHIN whose telephone number is (571)270-5764. The examiner can normally be reached Monday - Friday from 11:00AM to 7:00PM EST.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Said Broome can be reached at 571-272-2931. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/ANDREW SHIN/Examiner, Art Unit 2612
/MAURICE L. MCDOWELL, JR/Primary Examiner, Art Unit 2612