DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claim(s) 1-8, 10-18 and 20 are rejected under 35 U.S.C. 103 as being unpatentable over Shoss in view of Freeman et al (US 2019/0122404).
Claims 1 and 11, Shoss teaches a computing system and medium configured to perform video segmentation and processing operations, comprising: a memory device to store received video data; and processing circuitry configured to:
obtain video data from a video data source, the video data depicting a human user and an object in a scene;
Shoss: Fig. 6A and 6B: host 604 and objects 621, 622 and 685);
obtain context data from at least one other data source that provides non-video data, the context data related to an interaction of the human user with the object;
Shoss: Figs. 7A and 7B: a host individual picking up a product from a table), speech pertaining to the product (e.g., a host individual mentioning the product), eye gaze (e.g., a host individual looking at the product), and/or other activities, [0069-0070]);
analyze the context data to determine a shape of the object and a type of the interaction of the human user with the object, (Shoss: to identify objects…[0035]. Please see [0033-0038] for more detail on analysis of object and human interaction. Regarding context data which can be “news and weather information, sports highlights, product information, reviews of products and services, product promotion, educational material, how-to videos, advertising, and more, [0008] or icons, company logos, [0051, 0069, 0070, 0076]. Regarding shape which can be a “box” as in product 622 or can shape as in product 621 which is a subject matter keyword that the host selects and discusses) where the shape of the object is provided from a database of pre-trained object shapes (Shoss does not teach this feature, Freeman: [0080] An exemplary embodiment of a shape model training module 23 in the augmented reality system 1 will now be described in more detail with reference to FIG. 2, which shows the main elements of the shape model training module 23 as well as the data elements processed and generated by the shape model training module 23 for the trained shape models 15. As shown, the shape model training module 23 includes a shape model module 23a that retrieves training images 25a and corresponding user-defined feature points 25b from the training image database 25. The training image database 25 may store a plurality of training images 25a, each comprising the entire face of a respective person, including one or more facial features such as a mouth, eye or eyes, eyebrows, nose, chin, etc. For example, the training images 25a may include subject faces and facial features in different orientations and variations, such as front-on, slightly to one side, closed, pressed, open slightly, open wide, etc. The shape model training module 23 may include a face detector module 23b to detect and determine the location of a face in each retrieved training image 25a. The shape model module 23a generates and stores a global shape model 15a and a plurality of sub-shape models 15b for a trained shape model 15 in the model database 21, as will be described in more detail below. It will be appreciated that a plurality of trained shape models may be generated and stored in the model database 21, for example associated with respective different types of objects.
[0137] Returning to FIG. 15, at step S15-2, the initialised tracking module 3 receives captured image data from the camera 5, which can be an image in a sequence of images or video frames. At step S15-3, the tracking module determines if an object, a subject's face in this exemplary embodiment, was previously detected and located for tracking in a prior image or video frame. In subsequent iterations of the tracking process, the tracking module 3 may determine that the object was previously detected and located, for example from tracking data (not shown) stored by the system 1, the tracking data including a determined global object shape of the detected object, which can be used as the initialised global object shape for the current captured image. As this is the first time the tracking process is executed, processing proceeds to step S15-5 where the captured image data is processed by the object detector module 42 to detect an object in the image and to output a bounding box 51 of an approximate location for the detected object. At step S15-7, the tracking module 3 initialises the detected object shape using the trained global shape model 27, the statistics computed at step S8-11 above, and the corresponding global shape regression coefficient matrix 45 retrieved from the model database 21, based on the image data within the identified bounding box 51. FIG. 18A shows an example of an initialised object shape 71 within the bounding box 51, displayed over the captured image data 73. The trained shape model may be generated by the shape model training module 23 as described by the training process above. As shown, the candidate object shape at this stage is an initial approximation of the whole shape of the object within the bounding box 51, based on the global shape model 27. Accordingly, the location and shape of individual features of the object, such as the lips and chin in the example of FIG. 18A, are not accurate; and
generate a video stream that includes a virtual background overlaid on the video data, (Shoss: virtual background can replace actual background, [0009];
wherein the virtual background is modified to be segmented based on at least one outline of the human user, to remove the virtual background and cause the video stream to display portions of video data inside the at least one outline of the human user, and wherein the virtual background is modified to be further segmented based on (i) at least one segmentation outline [[of]] determined from the shape of the object retrieved from the database and (ii) the type of the interaction of the human user with the object, to remove the virtual background and cause the video stream to display other portions of the video data inside the at least one segmentation outline upon detection of the interaction of the human user with the objec
Shoss, via Figs. 7a and 7B, generates/composes a livestream video combining of the outline or contour of the host individual’s hand lifting up the objects/products 621 and 622 with the shape of a can or a box as the host individual interacts (i.e., discusses/demonstrates the products, [0066]. Furthermore, Shoss teaches removing a virtual background and replacing with a new background [0058]… and Freeman continues teaching on “segmentation outline upon detection of the interaction of the human user with the object. ([0130] The conditional probabilities are calculated by means of the statistics stored in the histogram building procedure employed as follows:
[00003]P(Cb,Cr|lip)=foregroundHistogram(Cb,Cr)numLipPixelsP(Cb,Cr|nonlip)=backgroundHistogram(Cb,Cr)numNonLipPixelsP(lip)=numLipPixelsnumTotalPixelsP(nonlip)=numNonLipPixelsnumTotalPixels. [0131] Once the probability map of being lip has been computed around the mouth area, the result will be used in order to reinforce the histogram quality through a clustering process which will produce a finer segmentation of the lip area. Please note examiner maps the interaction of the human user to the user’s application of makeup product on his/her face); and
generate a video stream that includes a virtual background overlaid on the video data, (Shoss: virtual background can replace actual background, [0009];
Therefore it would have been obvious to the ordinary artisan before the effective filing date to incorporate the teaching of Freeman into the teaching Shoss for the purpose of detailing to include at least a portion of a person's face including a region having a visible feature, retrieving augmentation values to augment said region of the captured image, computing at least one characteristic of the visible feature based on captured image data associated with the visible feature, modifying the retrieved augmentation values based on said computed at least one characteristic, augmenting pixel values in said region of the captured image based on the modified augmentation values, and outputting the captured image with the augmented pixel values for display.
Therefore, it would have been obvious to the ordinary artisan before the effective filing date to incorporate the teaching of Freeman into the teaching of Shoss for the purpose of providing helpful hints/tools in applying makeup products.
Claims 2 and 12, wherein the context data includes audio data with speech from the human user, and wherein the processing circuitry is further configured to: perform speech-to-text conversion of the audio data to produce text, (Shoss: speech-to-text, [0045, 0051] where the selection of the product can be based on information in an audio track 232. The information can include a combination of tones. The information can include utterances and/or speech from a host individual. The speech can be processed by a speech-to-text process for further analysis, [0045] where a product is selected and discussed by a host, [0011]) wherein the shape of the object is determined based on at least one keyword from the text)
Claims 3 and 13. The computing system of claim 2, wherein the processing circuitry is further configured to: identify the object in the video data based on the at least one keyword from the text. (Shoss: subject matter keyword, [0051]; …The selection of the product can be based on information in an audio track 232. The information can include a combination of tones. The information can include utterances and/or speech from a host individual. The speech can be processed by a speech-to-text process for further analysis, [0045]”.
Claims 4 and 14. The computing system of claim 2, wherein the shape of the object is provided from a database of pre-trained objects, and wherein a selection of the object from the database is performed using the at least one keyword from the text. (Freeman: The tracking module 3 also includes a visible feature detector 17 that automatically identifies regions of pixels in the captured image associated with one or more visible features of the detected face, such as predefined cheek, eye and lip regions of the person's face that have applied makeup products. [0076]).
Claims 5 and 15. The computing system of claim 2, further comprising:
a camera to capture the video data; and a microphone to capture the audio data. (Shoss: Some devices may include multiple cameras, including wide-angle, ultrawide, and telephoto lenses, (i.e., camera 608 of Fig. 5), along with stereo microphones, [0005] to provide information which can include a combination of tones. The information can include utterances and/or speech from a host individual. The speech can be processed by a speech-to-text process for further analysis. In embodiments, the defining the virtual background is based on the host individual's spoken words, [0045]).
Claims 6 and 16, further comprising a display device; wherein the processing circuitry is further configured to identify screen content output on the display device; and wherein the interaction of the human user with the object is ignored or identified based on the screen content. (Shoss: the identification of the foreground object as a product can include performing optical character recognition on text imprinted on a foreground object, and/or implementing other image recognition techniques. Further, the identification of the foreground object as a product can include scanning of an optical code such as a barcode that is imprinted on the product, [0026] or by motion of the product. For example, when a host individual picks up an object, the motion of the object can be detected and a virtual background can be defined and/or selected based on the motion of the object, [0027] on the screen of Fig. 7A and 7B).
Claims 7 and 17, further comprising a user input device; wherein the processing circuitry is further configured to identify a user input provided to the user input device, wherein the user input includes at least one of a keyboard, mouse, touch, or gesture input from the human user; and wherein the interaction of the human user with the object is ignored or identified based on the user input. (Shoss: The insertion point can be based on an absolute time, a time interval, a change of subject matter, motion of a foreground object, spoken words of a host individual, motion/gestures of a host individual, and/or other criteria, [0039, 0040, 0066, 0072]).
Claims 8 and 18, wherein the processing circuitry is further configured to: analyze the context data to determine a plurality of candidate objects in the scene for interaction; and select the object from the plurality of candidate objects based on at least one other interaction performed by the human user related to the object. (Shoss: Figs. 6A and 6B: products 621, 622, 685 where the individual interact with the products in Figs. 7A and 7B).
Claims 10 and 20, further comprising: communications circuitry to provide the video stream to another computing system in a video call or video conferencing session. (Shoss: the system 800 monitors text conversations in a chat window that is associated with a livestream video, [0083]).
Claim(s) 9 and 19 are rejected under 35 U.S.C. 103 as being unpatentable over Shoss in view of Freeman and further in view of in view of Lovemelt.
Claims 9 and 19, Shoss does not teach, “perform video post-processing on the video data based on the shape of the object and the type of the interaction of the human user”.
Lovemelt: [0085] In addition to previously described features and functionality, the video processing system 1020 is depicted as compositing functionality 1026 and effects functionality 1027. In an example, the compositing functionality 1026 is adapted to process camera video streams (e.g., camera feeds 760, 765) from a NIR/Visible camera system (e.g., dual camera system 300), and create a matte (e.g., luma matte 780) and generate output image and video (e.g., image with alpha channel 785) from the two respective video streams. The effects functionality 1027 is also adapted to implement post-processing video effects on all or a portion of the video streams (e.g., with the addition of additional video objects or layers, the distortion of colors, shapes, or perspectives in the video, and any other number of other video changes). In a further example, the video processing system 1020 may operate as a server, to receive and process video data obtained from the video capture system 1030, and to serve video data output to the video input/output system 1010.
Therefore, it would have been obvious to the ordinary artisan before the effective filing date to incorporate the teaching of Lovemelt into the teaching Shoss for the purpose of providing a post-processing effect in the composing technique to enhance communication in the immersive video environment.
Inquiry
Any inquiry concerning this communication or earlier communications from the examiner should be directed to PHUNG-HOANG J. NGUYEN whose telephone number is (571)270-1949. The examiner can normally be reached Reg. Sched. 6:00-3:00.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Duc Nguyen can be reached at 571-272-7503. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/PHUNG-HOANG J NGUYEN/Primary Examiner, Art Unit 2691