DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claim(s) 1-5, 8-12, 14-21, 23, and 29-30 are rejected under 35 U.S.C. 103 as being unpatentable over Choe et al. (US Pub 2013/0129307) in view of Zeng et al. (“Socratic Models: Composing Zero-Shot Multimodal Reasoning with Language”).
Regarding claim 1, Choe discloses a method for presenting sensor data, comprising:
at a computer system having one or more processors and memory (see fig. 1):
obtaining the sensor data from a plurality of sensor devices disposed in a physical environment during a time duration (para 0042, 0089; para 0061-0064);
generating one or more information items characterizing one or more signature events detected within the time duration in the sensor data (para 0064-066 – stored event records and generated textual description; also see fig. 3);
obtaining a natural language prompt (para 0081, 0084); and
in response to the natural language prompt:
presenting the multimodal output associated with the sensor data (para 0093- “may display information based on the textual description and overlay the information on or display the information beside a map or image of a geographical area, such that the information is visually associated with a particular geographical location. The information for display may be created, for example, at a server computer, and transmitted to a client computer for display”, 0095).
Choe does not disclose applying a large behavior model (LBM) to process the one or more information items and the natural language prompt jointly and generate a multimodal output associated with the sensor data; and
Zeng discloses applying a large behavior model (LBM) to process the one or more information items and the natural language prompt jointly (page 7, section 5.1 – describes extracting key moments from the video and captioning the frames and recursively summarizing into record of events and page 24-25, section D.3 – “(ii) Open-ended Q&A can be implemented by prompting the LM to complete the template: ‘{world-state history} Q: {question} A:’. We find that LMs such as GPT-3 can generate surprisingly meaningful results to binary yes or no questions, contextual reasoning questions, as well as temporal reasoning questions.” - ‘{world-state history} Q: {question} A’ means the prompt and items enter the model together) and generate a multimodal output associated with the sensor data (page 24-25, section D.3 – “Open-ended text prompts from a user, conditioned on an egocentric video, can yields three types of responses: a text-based response, a visual result, and/or an audio clip” and see fig. 11).
Therefore, it would have been obvious to a person of ordinary skilled in the art before the effective filing date of the claimed invention to modify Choe with the teachings of Zeng in order to “capture new multimodal functionalities zero-shot without having to rely on additional domain-specific data collection or model finetuning” (Zeng, page 9, section 6).
Regarding claim 2, Choe discloses wherein:
the sensor data is divided into a plurality of temporal windows, the method further comprising, each temporal window corresponding to at least a subset of sensor data (para 0073, 0078);
generating the one or more information items further includes, for each of a subset of temporal windows, processing the subset of sensor data to detect a respective signature event within each respective temporal window and generating a respective information item associated with the respective signature event (para 0064-0066); and
storing the one or more information items associated with the one or more signature events, the one or more information items including a timestamp and a location of each of the one or more signature events (para 0065; and also see para 0063).
Regarding claim 3, Choe discloses further comprising:
determining a behavior pattern based on the one or more signature events for the time duration of the sensor data (para 0064 – discloses complex events determined from combination of atomic events; 0068 – inferring traffic violation from sequence of underlying events);
generating a subset of the one or more information items describing the behavior pattern (para 0068 – generating report containing scene context description); and
providing the subset of the one or more information items of the behavior pattern associated with the sensor data (para 0068 – clickable event descriptions displayed in a browser).
Regarding claim 4, Zeng discloses applying the LBM further comprising:
providing, to the LBM, the natural language prompt and the one or more information items associated with one or more signature events (page 24, Section D.3 - the template: “{world-state history} Q: {question} A:” and see page 7, section 5.1 – “This is then passed as context to an LM to perform various reasoning tasks via text completion such as Q&A, for which LMs have demonstrated strong zero-shot performance”); and
in response to the natural language prompt, obtaining, from the LBM, the multimodal output describing the one or more signature events associated with the sensor data (page 24-25, section D.3 – “Open-ended text prompts from a user, conditioned on an egocentric video, can yields three types of responses: a text-based response, a visual result, and/or an audio clip” and see fig. 11).
Regarding claim 5, Zeng discloses wherein the natural language prompt includes a predefined mission, the predefined mission including a trigger condition (page 8, section 5.2 – discloses a prompt of expert chef which has trigger conditions in explicit if-then form).
Regarding claim 8, Choe discloses wherein the natural language prompt includes a user query entered on a user interface of an application executed on a client device, and the user query is received, in real time while or after the sensor data are collected (para 0081 and see figs. 8-9).
Regarding claim 9, the combination of Choe and Zeng discloses wherein the user query includes information defining the time duration, the method further comprising:
determining the time duration based on the user query (Choe, para 0082); and
extracting the one or more information items characterizing the sensor data for each temporal window that are included in the time duration Choe, (para 0083; also see 0065);
wherein the user query, the one or more information items in the time duration, and respective temporal timestamps are provided to the LBM (Zeng, page 7, section 5.1 and fig. 4 - world-state history carries timestamps as shown in fig. 4 ).
Regarding claim 10, Choe discloses wherein the user query includes information defining a location, the method further comprising:
selecting one of the plurality of sensor devices based on the user query (para 0086 – “As a result of the selection, a set of events are displayed, including the selected event, and other events that precede and/or that follow the event at the same location or from the same video camera”);
identifying a subset of sensor data captured the selected one of the plurality of sensor devices (see page 10, claim 39 – “capturing a plurality of video sequences from a plurality of respective video sources, each video sequence including a plurality of video images;
for each video sequence, automatically detecting one or more events that involve at least two agents in the video sequence, each event associated with a set of video images;”); and
extracting the one or more information items characterizing the sensor data associated with the selected one of the plurality of sensor devices (para 0083 – steps 803-804 discloses the events matching the search criteria are retrieved and returned with their description with geographic and time information; also see fig. 10E).
Regarding claim 11, Choe discloses wherein the user query includes information defining a location, the method further comprising:
identifying a region of interest corresponding to the location in the sensor data captured by a first sensor (para 0075-0077); and
extracting the one or more information items characterizing the sensor data associated with the region of interest (para 0075-0078 – “a user may wish to search for all passenger pickups by vehicles on a particular block within a given time period”).
Regarding claim 12, see rejection of claim 9.
Regarding claim 14, Choe discloses wherein the method is implemented by a server system, and the server system is coupled to a client device that executes an application, the method further comprising:
enabling display of a user interface on the application, including receiving the natural language prompt via the user interface and providing the multimodal output characterizing the sensor data (para 0093- “may display information based on the textual description and overlay the information on or display the information beside a map or image of a geographical area, such that the information is visually associated with a particular geographical location. The information for display may be created, for example, at a server computer, and transmitted to a client computer for display”, 0095; also see figs. 7, 9, 10C)
Regarding claim 15, Zeng discloses wherein the natural language prompt defines a reply language, and the multimodal output is provided by the LBM in the reply language (fig. 11 and see page 8, section 5.2 prompt – instructions on how to reply if-then format).
Regarding claim 16, Choe discloses wherein the multimodal output includes one or more of: description, timestamp, numeral information, statistic summary, warning message , and recommended action associated with one or more signature events (para 0066, 0077).
Regarding claim 17, Choe discloses wherein the multimodal output includes one or more of: textual statements, software code, an image or video, an information dashboard having a predefined format, a user interface, and a heatmap (see figs. 6 and 11B).
Regarding claim 18, Choe discloses wherein for a temporal window corresponding to a subset of sensor data, the method further comprising:
using at least an event projection model to detect one or more signature events based on the subset of sensor data within the temporal window (para 0052 -0055 – “graph grammar module 230 models the content of video images in terms of objects in a scene, scene elements, and their relations. The model defines the visual vocabulary, attributes of scene elements, and their production rule”).
Regarding claim 19, Choe discloses further comprising:
storing the one or more information items and/or the multimodal output in a database, in place of the sensor data measured by the plurality of sensor devices (para 0065 and see fig. 8 – searching based on events not video data).
Regarding claim 20, Choe discloses further comprising:
processing the sensor data to generate one or more sets of intermediate items successively and iteratively, until generating the one or more information items (para 0050-0052; para 0064 – “events may include atomic events (e.g., events that cannot be further broken down)”).
Regarding claim 21, Zeng discloses wherein the LBM includes a large language model (LLM) (page 19, fig. 7 – “large language models (LMs, e.g., GPT-3, RoBERTa)”).
Regarding claim 23, Choe discloses wherein the sensor data include video data streamed by cameras that are disposed at a venue (para 0042), and the multimodal output includes a chart indicating a plurality of site states or snapshots associated with respective feature events (para 0084 and fig. 9).
Regarding claims 29 and 30, see rejection of claim 1.
Claim 26 is rejected under 35 U.S.C. 103 as being unpatentable over Choe et al. (US Pub 2013/0129307) in view of Zeng et al. (“Socratic Models: Composing Zero-Shot Multimodal Reasoning with Language”) and in further view of deSa et al (US Pub 2021/0057101).
Regarding claim 26, Choe in view of Zeng discloses the method of claim 1.
Choe in view of Zeng does not disclose wherein the sensor data are provided by a radar disposed in a room, and the multimodal output includes an avatar that is enabled for display in accordance with a determination that the radar detects a presence of a person in the room.
deSa discloses wherein the sensor data are provided by a radar disposed in a room, and the multimodal output includes an avatar that is enabled for display in accordance with a determination that the radar detects a presence of a person in the room (para 0034, para 0040).
Therefore, it would have been obvious to a person of ordinary skilled in the art before the effective filing date of the claimed invention to modify Choe in view of Zeng with the teachings of deSa in order to track location of a user with high accuracy (deSa, para 0035).
Allowable Subject Matter
Claims 13 and 22 are objected to as being dependent upon a rejected base claim, but would be allowable if rewritten in independent form including all of the limitations of the base claim and any intervening claims.
Conclusion
Any inquiry concerning this communication or earlier communications from the examiner should be directed to NAFIZ E HOQUE whose telephone number is (571)270-1811. The examiner can normally be reached M-F 8-5.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Ahmad Matar can be reached at (571)272-7488. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/NAFIZ E HOQUE/ Primary Examiner, Art Unit 2693