DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Priority
2. Receipt is acknowledged of papers submitted under 35 U.S.C. 119(a)-(d), which papers have been placed of record in the file.
Information Disclosure Statement
3. The information disclosure statement (IDS) submitted on 02/12/2025, 09/04/2025 and 08/10/2026. The submission is in compliance with the provisions of 37 CFR 1.97. Accordingly, the information disclosure statement is being considered by the examiner.
Claim Rejections - 35 USC § 103
4. In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
5. The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
6. The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows:
1. Determining the scope and contents of the prior art.
2. Ascertaining the differences between the prior art and the claims at issue.
3. Resolving the level of ordinary skill in the pertinent art.
4. Considering objective evidence present in the application indicating obviousness or nonobviousness.
7. Claim(s) 1-11 and 15-20 is/are rejected under 35 U.S.C. 103 as being unpatentable over Goodrich et al. (US 2021/0065448 A1) in view of Narasimha et al. (US 2015/0243035 A1).
8. With reference to claim 1, Goodrich teaches A method comprising: receiving an image depicting an object; (“FIG. 9 is a flowchart illustrating a method 900 of performing conversion passes for processing the image and depth data which may be performed in conjunction with the method 800 for generating a 3D message as described above, according to some example embodiments. The method 900 may be embodied in computer-readable instructions for execution by one or more computer processors such that the operations of the method 900 may be performed in part or in whole by the messaging client application 104, particularly with respect to respective components of the annotation system 206 described above in FIG. 7; … the image and depth data receiving module 702 receives image data and depth data captured by an optical sensor (e.g., camera) of the client device 102. In an example, the depth data includes a depth map corresponding to the image data. … using face tracking and portrait segmentation to generate a depth map of a person.” [0154-0156] “the biometric components 2330 may include components to detect expressions (e.g., hand expressions, facial expressions, vocal expressions, body gestures, or eye tracking), measure biosignals (e.g., blood pressure, heart rate, body temperature, perspiration, or brain waves), identify a person (e.g., voice identification, retinal identification, facial identification, fingerprint identification, or electroencephalogram based identification), and the like.” [0247]) Goodrich also teaches the object depicted in the image by applying one or more machine learning models to predict depths of pixels corresponding to the object; (“the machine learning model can be a deep neural network or a convolutional neural network that provides a prediction of depth data based on the captured image data, and the machine learning model receives the captured image data as an input, and generates a depth map as an output. In some implementations, the machine learning model executes on a neural network processor or a graphics processing unit of the client device.” [0168] “machine learning models can be applied in a beautification operation such as convolutional neural networks, generative adversarial networks, and the like. Such machine learning models can be utilized to preserve facial feature structures, smooth blemishes or remove wrinkles, or preserve facial skin texture in facial image data.” [0181]) Goodrich further teaches displaying an augmented reality (AR) liquid element approaching the object from a direction in the image; (“a view 1300 is provided for display on a display of a client device (e.g., the client device 102), The view 1300 includes an image of a representation of a user's portrait (e.g., including a face). Selectable graphical element 1305 is provided for display in the view 1300. In an embodiment, selectable graphical element 1305 corresponds to an augmented reality content generator for generating a 3D message and applying 3D effects and other image processing operations as discussed further herein. … This display in the view 1350 can be updated to render the 3D effects associated with the 3D message in response to receiving sensor data (e.g., movement data, gyroscopic sensor data, and the like) in which the user is moving the client device. In an example, depending on the relative position of the client device with respect to a viewing user, the 3D effects can be updated for presentation on the display of the client device taking into account the change in position. For example, if the display of the client device is tilted in a particular manner to a first position, one set of 3D effects may be rendered and provided for display, and when the client device is moved to a different position, a second set of 3D effects may be rendered to update the image and indicate a change in viewing perspective, which provides a more 3D viewing experience to the viewing user.” [0202-0203] “The output components 2326 may include visual components (e.g., a display such as a plasma display panel (PDP), a light emitting diode (LED) display, a liquid crystal display (LCD), a projector, or a cathode ray tube (CRT)), acoustic components (e.g., speakers), haptic components (e.g., a vibratory motor, resistance mechanisms), other signal generators, and so forth.” [0246])
PNG
media_image1.png
795
537
media_image1.png
Greyscale
Goodrich does not explicitly teach generating a dense depth reconstruction of the object, determining that a depth of the AR liquid element matches a depth of a portion of the object based on the dense depth reconstruction; and modifying, in response to the determining, at least one of a trajectory of the AR liquid element, or a color of the portion of the object that matches the depth of the AR liquid element. This is what Narasimha teaches (“If the 3D features are point features, different kinds of iterative closed point (ICP) methods may be employed to determine the N-th accurate transformation by matching at least part of the N-th plurality of 3D features with at least part of the 3D features of the 3D model. The ICP method requires an initial guess. The estimated N-th coarse transformation may be used as an initial guess for the ICP method to match between the at least part of the N-th plurality of 3D features and the at least part of the 3D features of the 3D model. … In step 1009 of determining the N-th accurate transformation, points (i.e. 3D features) from different rigid parts (e.g. nose and cheek of the head) may be treated differently in ICP when aligning the N-th plurality of 3D features with the 3D features of the 3D model. … A Kalman filter may be used to smooth the determined N-th accurate transformation. … The FGM framework comprises dynamically updating the model with live depth measurements while using it to register incoming depth images. Particularly, the 3D model (bump Image in the present case) is initialized with the first input depth image and then augmented by the subsequent input depth images. For the N-th input depth image, a current projected depth image may be generated from the current 3D model by transforming the 3D model with a previously estimated pose (i.e. the coarse or accurate transformation of a (N-1)-th input depth image) or with the N-th coarse transformation and rendering the transformed 3D model in OpenGL context, and then the N-th depth image is aligned to the current projected depth image.” [0074-0078] “when a random forest is to be used to estimate the coarse pose from an input image (e.g. image 2003 or 2005), patches are extracted (either at random or in a dense sampling scheme) from the input image using face detection. Then, the patches are propagated through the trained pose model (i.e. a trained forest of binary trees in this example). When the patches reach leaf nodes, the trained mean and variance values for profile angle values at these leaves are used to estimate the coarse pose for the observed face image.” [0105] “In Augmented Reality (AR) applications, virtual visual content (like a computer generated object) may be overlaid onto an image of an object of interest based on a reconstructed 3D model of the object of interest. In one example of AR applications, a virtual glasses may be generated and overlaid onto an image of a human head. A 3D model of the human head may be required to select a proper size of the virtual glasses or adjust the shape of the virtual glasses. Depth images of the head may be captured by using a depth camera. The 3D model of the head could be generated according to the method disclosed in this invention. The virtual glasses could be overlaid onto any of the captured depth images of the head according to at least part of the reconstructed 3D model of the head.” [0111] “A depth image particularly is a 2D image with a corresponding depth map.” [0118]) Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the teachings of Narasimha into Goodrich, in order to have less constraints about displacement between any input images.
9. With reference to claim 2, Goodrich does not explicitly teach first and second stages of the one or more machine learning models are part of a common portion of the one or more machine learning models that simultaneously output a depth of a point of interest and relative depths of each pixel in the portion of the image. This is what Narasimha teaches (“The transformation 2007 (i.e. the first coarse transformation) may be determined by using a machine learning method according to a trained pose model and at least part of the input image.” [0058] “Step 1005 provides an N-th (e.g. a first (FIG. 1A) or second (FIG. 1B)) input depth image of the object of interest, wherein an N-th image COS is associated with the N-th input depth image.” [0065] “the transformation 2008 (indicated by dash lines 2008 in FIG. 2) between the face COS 2002 and the image COS 2006 is determined as the N-th coarse transformation. In this example, the transformation 2008 describes a pose of the depth camera 2010 relative to the head or face 2001 when the depth camera 2010 captures the depth image 2005. … The transformation 2008 (i.e. the N-th coarse transformation) may be determined by using a machine learning method according to a trained pose model and at least part of the input image.” [0069-0070] “To perform temporal mean between the two depth images (N-th input depth image and aligned current projected depth image), it is essential to weigh the depth from the current projected depth image based on the weighted frequency of its appearance, i.e. its corresponding confidence value. The word "weighted" is stressed as every new depth pixel in the N-th input depth image can have a weight of at most 1. Since the depth precision changes with distance it is better to encode the uncertainty of depth at different distances in the weights. This is done by weighing each new depth pixel in the N-th input depth image” [0085] “The updating of Bump Image depends on the confidence values associated with each pixel in the Bump Image and the newly estimated confidence values.” [0087] “A 3D model may describe a geometry for an object or a generic geometry for a group of objects. For example, a 3D model may be specific for an object. A 3D model may not be specific for an object, but may describe a generic geometry for a group of similar objects. A similar object may belong to the same type of object and share some common properties. For example, faces of different people are of same type since they are a respective face that has eye, mouth, ear, nose, etc.” [0137]) Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the teachings of Narasimha into Goodrich, in order to have less constraints about displacement between any input images.
10. With reference to claim 3, Goodrich teaches the first stage comprises a first machine learning model (“the machine learning model can be a deep neural network or a convolutional neural network that provides a prediction of depth data based on the captured image data, and the machine learning model receives the captured image data as an input, and generates a depth map as an output. In some implementations, the machine learning model executes on a neural network processor or a graphics processing unit of the client device.” [0168] “machine learning models can be applied in a beautification operation such as convolutional neural networks, generative adversarial networks, and the like. Such machine learning models can be utilized to preserve facial feature structures, smooth blemishes or remove wrinkles, or preserve facial skin texture in facial image data.” [0181])
Goodrich does not explicitly teach the second machine learning model stage comprises a second machine learning model. This is what Narasimha teaches (“the transformation 2008 (indicated by dash lines 2008 in FIG. 2) between the face COS 2002 and the image COS 2006 is determined as the N-th coarse transformation. In this example, the transformation 2008 describes a pose of the depth camera 2010 relative to the head or face 2001 when the depth camera 2010 captures the depth image 2005. … The transformation 2008 (i.e. the N-th coarse transformation) may be determined by using a machine learning method according to a trained pose model and at least part of the input image.” [0069-0070] “In Augmented Reality (AR) applications, virtual visual content (like a computer generated object) may be overlaid onto an image of an object of interest based on a reconstructed 3D model of the object of interest.” [0111]) Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the teachings of Narasimha into Goodrich, in order to have less constraints about displacement between any input images.
11. With reference to claim 4, Goodrich teaches an output of the first stage comprises a distance between a camera used to capture the image and the point of interest in a first coordinate system global to the image. (“Such landmark identification procedures may be used for any such objects. In some embodiments, a set of landmarks forms a shape. Shapes can be represented as vectors using the coordinates of the points in the shape. One shape is aligned to another with a similarity transform (allowing translation, scaling, and rotation) that minimizes the average Euclidean distance between shape points. The mean shape is the mean of the aligned training shapes. In some embodiments, a search for landmarks from the mean shape aligned to the position and size of the face determined by a global face detector is started. Such a search then repeats the steps of suggesting a tentative shape by adjusting the locations of shape points by template matching of the image texture around each point and then conforming the tentative shape to a global shape model until convergence occurs.” [0060-0061] “The image and depth data receiving module 702 receives images and depth data captured by a client device 102. For example, an image is a photograph captured by an optical sensor (e.g., camera) of the client device 102. An image includes one or more real-world features, such as a user's face or real-world object(s) detected in the image. In some embodiments, an image includes metadata describing the image. For example, the depth data includes data corresponding to a depth map including depth information based on light rays emitted from a light emitting module directed to an object (e.g., a user's face) having features with different depths (e.g., eyes, ears, nose, lips, etc.). By way of example, a depth map is similar to an image but instead of each pixel providing a color, the depth map indicates distance from a camera to that part of the image (e.g., in absolute terms, or relative to other pixels in the depth map).” [0132] “the machine learning model can be a deep neural network or a convolutional neural network that provides a prediction of depth data based on the captured image data, and the machine learning model receives the captured image data as an input, and generates a depth map as an output. In some implementations, the machine learning model executes on a neural network processor or a graphics processing unit of the client device.” [0168])
12. With reference to claim 5, Goodrich teaches an output of the second stage comprises a depth from the point of interest for each foreground pixel and a segmentation mask for each pixel in the portion of the image, the depth from the point of interest for each foreground pixel in the portion of the image being represented in a second coordinate system local to the portion of the image. (“the image and depth data processing module 706 generates a segmentation mask based at least in part on the image data. In an embodiment, the image and depth data processing module 706 determines the segmentation mask using a convolutional neural network to perform dense prediction tasks where a prediction is made for every pixel to assign the pixel to a particular object class (e.g., face/portrait or background), and the segmentation mask is determined based on the groupings of the classified pixels (e.g., face/portrait or background).” [0161] “the image and depth data processing module 706 provides a post-process foreground image by applying, using the depth normal map, a 3D effect(s) to a foreground region of the image data. At operation 916, the rendering module 710 generates a view of the 3D message using at least the background inpainted image, the inpainted depth map, and the post-processed foreground image, which are assets that are included the generated 3D message.” [0165-0166])
13. With reference to claim 6, Goodrich does not explicitly teach generating a dense point cloud based on the dense depth reconstruction. This is what Narasimha teaches (“Step 1006 provides an N-th plurality of 3D features in the N-th image COS according to the Nth input depth image in a similar way as in step 1002b. In the example shown in FIG. 2, the N-th plurality of 3D features at least contains features of the human head or face 2001. For example, the N-th plurality of 3D features is a point cloud comprising 3D points of the head surface. 3D features (e.g. 3D points) of at least one rigid part of the object of interest may be preferred to be determined to be as at least part of the N-th plurality of 3D features.” [0066] “when a random forest is to be used to estimate the coarse pose from an input image (e.g. image 2003 or 2005), patches are extracted (either at random or in a dense sampling scheme) from the input image using face detection. Then, the patches are propagated through the trained pose model (i.e. a trained forest of binary trees in this example). When the patches reach leaf nodes, the trained mean and variance values for profile angle values at these leaves are used to estimate the coarse pose for the observed face image.” [0105] “In Augmented Reality (AR) applications, virtual visual content (like a computer generated object) may be overlaid onto an image of an object of interest based on a reconstructed 3D model of the object of interest. In one example of AR applications, a virtual glasses may be generated and overlaid onto an image of a human head. A 3D model of the human head may be required to select a proper size of the virtual glasses or adjust the shape of the virtual glasses. Depth images of the head may be captured by using a depth camera. The 3D model of the head could be generated according to the method disclosed in this invention. The virtual glasses could be overlaid onto any of the captured depth images of the head according to at least part of the reconstructed 3D model of the head.” [0111]) Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the teachings of Narasimha into Goodrich, in order to have less constraints about displacement between any input images.
14. With reference to claim 7, Goodrich does not explicitly teach applying the AR liquid element based on at least one of a geometry of a body of a person, hair of the person, clothing of the person, or one or more accessories worn by the person. This is what Narasimha teaches (“Step 1006 provides an N-th plurality of 3D features in the N-th image COS according to the Nth input depth image in a similar way as in step 1002b. In the example shown in FIG. 2, the N-th plurality of 3D features at least contains features of the human head or face 2001. For example, the N-th plurality of 3D features is a point cloud comprising 3D points of the head surface. 3D features (e.g. 3D points) of at least one rigid part of the object of interest may be preferred to be determined to be as at least part of the N-th plurality of 3D features. For example, rigid parts of the head may be nose, cheek, and ear. Points of the nose may be selected as at least part of the N-th plurality of 3D features. Points of the cheek may be selected as at least part of the N-th plurality of 3D features. Points of the nose and points of the cheek may be comprised in the 3D model.” [0066] “In Augmented Reality (AR) applications, virtual visual content (like a computer generated object) may be overlaid onto an image of an object of interest based on a reconstructed 3D model of the object of interest. In one example of AR applications, a virtual glasses may be generated and overlaid onto an image of a human head. A 3D model of the human head may be required to select a proper size of the virtual glasses or adjust the shape of the virtual glasses. Depth images of the head may be captured by using a depth camera. The 3D model of the head could be generated according to the method disclosed in this invention. The virtual glasses could be overlaid onto any of the captured depth images of the head according to at least part of the reconstructed 3D model of the head. An embodiment of determining poses of the head and a 3D model of the head may be used in AR shopping applications, e.g. shopping glasses or hat. Particularly, as deformable parts of the head (mouth) may not be selected as 3D features, the accuracy of pose estimation and 3D reconstruction would be improved for upper parts (rigid parts, e.g. nose, cheek) of the head.” [0111]) Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the teachings of Narasimha into Goodrich, in order to have less constraints about displacement between any input images.
15. With reference to claim 8, Goodrich teaches displaying the AR liquid element on a first portion of a person depicted in a first frame of a video, wherein the person is positioned at a first location in the first frame; (“when a particular modification is selected along with content to be transformed, elements to be transformed are identified by the computing device, and then detected and tracked if they are present in the frames of the video. The elements of the object are modified according to the request for modification, thus transforming the frames of the video stream. Transformation of frames of a video stream can be performed by different methods for different kinds of transformation.” [0057] “other methods and algorithms suitable for face detection can be used. For example, in some embodiments, features are located using a landmark which represents a distinguishable point present in most of the images under consideration. For facial landmarks, for example, the location of the left eye pupil may be used. In an initial landmark is not identifiable (e.g., if a person has an eyepatch), secondary landmarks may be used.” [0060] “a computer animation model to transform image data can be used by a system where a user may capture an image or video stream of the user (e.g., a selfie) using a client device 102 having a neural network operating as part of a messaging client application 104 operating on the client device 102. The transform system operating within the messaging client application 104 determines the presence of a face within the image or video stream and provides modification icons associated with a computer animation model to transform image data, or the computer animation model can be present as associated with an interface described herein. The modification icons include changes which may be the basis for modifying the user's face within the image or video stream as part of the modification operation. Once a modification icon is selected, the transform system initiates a process to convert the image of the user to reflect the selected modification icon (e.g., generate a smiling face on the user). In some embodiments, a modified image or video stream may be presented in a graphical user interface displayed on the mobile client device as soon as the image or video stream is captured and a specified modification is selected. The transform system may implement a complex convolutional neural network on a portion of the image or video stream to generate and apply the selected modification. That is, the user may capture the image or video stream and be presented with a modified result in real time or near real time once a modification icon has been selected.” [0063] “The output components 2326 may include visual components (e.g., a display such as a plasma display panel (PDP), a light emitting diode (LED) display, a liquid crystal display (LCD), a projector, or a cathode ray tube (CRT)), acoustic components (e.g., speakers), haptic components (e.g., a vibratory motor, resistance mechanisms), other signal generators, and so forth.” [0246]) Goodrich also teaches determining that the person has moved from the first location to a second location in a second frame of the video; and updating a display position of the AR liquid element in the second frame to maintain the display of the AR liquid element on data representing the person depicted in the image. (“the sensor data receiving module 704 receives movement data from a movement sensor (e.g., gyroscope, motion sensor, touchscreen, etc.). In an embodiment, messaging client application 104 receives sensor data captured by a sensor of the client device 102, such as a location or movement sensor.” [0187] “the rendering module 710 renders the updated view of the 3D message. The updated view of the 3D is provided for display on a display of the client device the client device 102). FIG. 12 illustrates example user interfaces depicting a carousel for selecting and applying an augmented reality content generator to media content (e.g., an image or video), and presenting the applied augmented reality content generator in the messaging client application 104 (or the messaging system 100), according to some embodiments.” [0190-0191] “a disparity map is generated based at least in part on a distance between a first pixel from a first image captured by the first camera and a second pixel from a second image captured by the second camera, the first pixel and second pixel corresponding to a same object.” [0199] “This display in the view 1350 can be updated to render the 3D effects associated with the 3D message in response to receiving sensor data (e.g., movement data, gyroscopic sensor data, and the like) in which the user is moving the client device. In an example, depending on the relative position of the client device with respect to a viewing user, the 3D effects can be updated for presentation on the display of the client device taking into account the change in position. For example, if the display of the client device is tilted in a particular manner to a first position, one set of 3D effects may be rendered and provided for display, and when the client device is moved to a different position, a second set of 3D effects may be rendered to update the image and indicate a change in viewing perspective, which provides a more 3D viewing experience to the viewing user.” [0203] “The output components 2326 may include visual components (e.g., a display such as a plasma display panel (PDP), a light emitting diode (LED) display, a liquid crystal display (LCD), a projector, or a cathode ray tube (CRT)), acoustic components (e.g., speakers), haptic components (e.g., a vibratory motor, resistance mechanisms), other signal generators, and so forth.” [0246])
16. With reference to claim 9, Goodrich teaches replacing data representing the depiction of a person with one or more visual effects. (“augmented reality content generators, augmented reality content items, overlays, image transformations, AR images and similar terms refer to modifications that may be made to videos or images. This includes real-time modification which modifies an image as it is captured using a device sensor and then displayed on a screen of the device with the modifications. … Data and various systems using augmented reality content generators or other such transform systems to modify content using this data can thus involve detection of objects (e.g., faces, hands, bodies, cats, dogs, surfaces, objects, etc.), tracking of such objects as they leave, enter, and move around the field of view in video frames, and the modification or transformation of such objects as they are tracked.” [0054-0055] “Such modifications may involve changing color of areas; removing at least some part of areas from the frames of the video stream; including one or more new objects into areas which are based on a request for modification; and modifying or distorting the elements of an area or object. In various embodiments, any combination of such modifications or other similar modifications may be used.” [0058] “a 3D augmented reality content generator refers to a real-time special effect and/or sound that may be added to a 3D message and modifies image and/or depth data.” [0099] “such beautification techniques can modify facial image data in the digital domain, such as slimming cheeks, enlarging eyes, smoothing skin, brightening teeth or skin, removing blemishes or wrinkles, changing eye color, shrinking sagging skin, enhancing skin color, adding facial tattoos or markings, and the like.” [0174])
17. With reference to claim 10, Goodrich teaches associating the AR liquid element with a first distance between a camera used to capture the image and the AR liquid element. (“An image includes one or more real-world features, such as a user's face or real-world object(s) detected in the image. In some embodiments, an image includes metadata describing the image. For example, the depth data includes data corresponding to a depth map including depth information based on light rays emitted from a light emitting module directed to an object (e.g., a user's face) having features with different depths (e.g., eyes, ears, nose, lips, etc.). By way of example, a depth map is similar to an image but instead of each pixel providing a color, the depth map indicates distance from a camera to that part of the image (e.g., in absolute terms, or relative to other pixels in the depth map).” [0132] “The output components 2326 may include visual components (e.g., a display such as a plasma display panel (PDP), a light emitting diode (LED) display, a liquid crystal display (LCD), a projector, or a cathode ray tube (CRT)), acoustic components (e.g., speakers), haptic components (e.g., a vibratory motor, resistance mechanisms), other signal generators, and so forth.” [0246])
18. With reference to claim 11, Goodrich teaches in response to determining that a second distance is greater than the first distance, applying a first visual effect to the portion of the image; and in response to determining that the second distance is less than the first distance, applying a second visual effect to the portion of the image. (“a disparity map is generated based at least in part on a distance between a first pixel from a first image captured by the first camera and a second pixel from a second image captured by the second camera, the first pixel and second pixel corresponding to a same object. The disparity map is an image where each pixel includes a distance value between a pixel from the first image to corresponding pixel from the second image. First pixels of a first object in the disparity map have a greater brightness than second pixels of a second object in the disparity map, the first pixels having a lesser depth values than second depth values of the second pixels. … renders the 3D effects for display as specified in the received message. Further, this other user can provide movement to the receiving client device, which in response, initiates a re-rendering of the 3D effects in which the perspective of the scene that is being viewed by the viewer is changed based on the provided movement.” [0199-0200] “This display in the view 1350 can be updated to render the 3D effects associated with the 3D message in response to receiving sensor data (e.g., movement data, gyroscopic sensor data, and the like) in which the user is moving the client device. In an example, depending on the relative position of the client device with respect to a viewing user, the 3D effects can be updated for presentation on the display of the client device taking into account the change in position. For example, if the display of the client device is tilted in a particular manner to a first position, one set of 3D effects may be rendered and provided for display, and when the client device is moved to a different position, a second set of 3D effects may be rendered to update the image and indicate a change in viewing perspective, which provides a more 3D viewing experience to the viewing user.” [0203])
19. With reference to claim 15, Goodrich teaches the one or more machine learning models comprise a neural network, the neural network being trained to establish a relationship between image portions depicting different orientations of human bodies and depths of points of interest of the human bodies. (“Real-time video processing can be performed with any kind of video data (e.g., video streams, video files, etc.) saved in a memory of a computerized system of any kind. For example, a user can load video files and save them in a memory of a device, or can generate a video stream using sensors of the device. Additionally, any objects can be processed using a computer animation model, such as a human's face and parts of a human body, animals, or non-living things such as chairs, cars, or other objects.” [0056] “a computer animation model to transform image data can be used by a system where a user may capture an image or video stream of the user (e.g., a selfie) using a client device 102 having a neural network operating as part of a messaging client application 104 operating on the client device 102. The transform system operating within the messaging client application 104 determines the presence of a face within the image or video stream and provides modification icons associated with a computer animation model to transform image data, or the computer animation model can be present as associated with an interface described herein. The modification icons include changes which may be the basis for modifying the user's face within the image or video stream as part of the modification operation. Once a modification icon is selected, the transform system initiates a process to convert the image of the user to reflect the selected modification icon (e.g., generate a smiling face on the user). In some embodiments, a modified image or video stream may be presented in a graphical user interface displayed on the mobile client device as soon as the image or video stream is captured and a specified modification is selected. The transform system may implement a complex convolutional neural network on a portion of the image or video stream to generate and apply the selected modification.” [0063] “the machine learning model can be a deep neural network or a convolutional neural network that provides a prediction of depth data based on the captured image data, and the machine learning model receives the captured image data as an input, and generates a depth map as an output. In some implementations, the machine learning model executes on a neural network processor or a graphics processing unit of the client device.” [0168])
Goodrich does not explicitly teach orientations of human bodies. This is what Narasimha teaches (“The head or face orientation from ground truth data is obtained by using a marker based tracking method that uses a known marker pose with respect to the camera coordinate system. Positive patches are extracted from facial region and negative patches are extracted from non-facial region. Each positive patch is annotated with a vector v=(v.sub.x,v.sub.y) that joins the center of the patch to the nose tip and the head orientation .theta.=(.theta..sub.yaw, .theta..sub.pitch, .theta..sub.rol1). A number of such positive patches are extracted from each depth image. For negative patches however, there is no associated vector v and orientation .theta.. These extracted positive and negative patches are then used to train the Random Forests algorithm.” [0098]) Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the teachings of Narasimha into Goodrich, in order to have less constraints about displacement between any input images.
20. With reference to claim 16, Goodrich teaches training the one or more machine learning models by performing operations comprising: receiving a plurality of training data sets, each of the plurality of training data sets comprising a training portion representing a training person depicted in an image and a corresponding depth data; (“Real-time video processing can be performed with any kind of video data (e.g., video streams, video files, etc.) saved in a memory of a computerized system of any kind. For example, a user can load video files and save them in a memory of a device, or can generate a video stream using sensors of the device. Additionally, any objects can be processed using a computer animation model, such as a human's face and parts of a human body, animals, or non-living things such as chairs, cars, or other objects.” [0056] “machine learning techniques and heuristics are utilized to generate depth maps in instances in which a given client device does not include appropriate hardware (e.g., depth sensing camera) that enables capture depth information. Such machine learning techniques can train a machine learning model from training data from shared 3D messages in an example (or other image data), Such heuristics include using face tracking and portrait segmentation to generate a depth map of a person.” [0156] “the machine learning model can be a deep neural network or a convolutional neural network that provides a prediction of depth data based on the captured image data, and the machine learning model receives the captured image data as an input, and generates a depth map as an output. In some implementations, the machine learning model executes on a neural network processor or a graphics processing unit of the client device.” [0168] “machine learning models can be applied in a beautification operation such as convolutional neural networks, generative adversarial networks, and the like. Such machine learning models can be utilized to preserve facial feature structures, smooth blemishes or remove wrinkles, or preserve facial skin texture in facial image data.” [0181]) Goodrich also teaches applying the one or more machine learning models to a first training portion of a first training data set to predict an estimated depth data for a given point of interest of the training person; (“the image and depth data processing module 706 generates a segmentation mask based at least in part on the image data. In an embodiment, the image and depth data processing module 706 determines the segmentation mask using a convolutional neural network to perform dense prediction tasks where a prediction is made for every pixel to assign the pixel to a particular object class (e.g., face/portrait or background), and the segmentation mask is determined based on the groupings of the classified pixels (e.g., face/portrait or background).” [0161] “the machine learning model can be a deep neural network or a convolutional neural network that provides a prediction of depth data based on the captured image data, and the machine learning model receives the captured image data as an input, and generates a depth map as an output. In some implementations, the machine learning model executes on a neural network processor or a graphics processing unit of the client device.” [0168])
Goodrich does not explicitly teach computing a deviation between the estimated depth data and the ground-truth depth data associated with the first training portion; and updating one or more parameters of the one or more machine learning models based on the computed deviation. This is what Narasimha teaches (“We employ the running average to integrate new measurements (i.e. a plurality of 3D features determined from the N-th input depth image) in the 3D model while reducing input noise. In order to minimize the noise a temporal mean filter is employed and points lying within, e.g., 1 cm deviation to the 3D model are subjected to mean filtering. In order to perform temporal mean filtering, another buffer with similar dimensions of the Bump Image is maintained, this buffer/image also known as confidence mask has one-to-one correspondence to all the pixels in the Bump Image and it records the weighted frequency of appearance of each pixel in the Bump Image.” [0082] “A Kalman filter is a 2-step filtering process that maintains a state for the object and uses the observations from the data to update the state. The first step is to predict the state in the current frame based on the state in the previous frame. The second step is to update the predicted state by taking into account the observations in the current frame. In our system, the nose tip location (x,y,z values) and velocity of the nose tip (along x,y,z directions) are maintained as the state. The observations are the predicted nose tip location from the random forest. In frames where the random forest returns a reliable nose tip estimate, we perform Kalman prediction and update steps to obtain the filtered nose tip location. In frames where the random forest does not return a reliable nose tip estimate, only the Kalman prediction step is performed. This allows the Kalman filter to continuously track and smooth the face center.” [0090] “The head or face orientation from ground truth data is obtained by using a marker based tracking method that uses a known marker pose with respect to the camera coordinate system. Positive patches are extracted from facial region and negative patches are extracted from non-facial region. Each positive patch is annotated with a vector v=(v.sub.x,v.sub.y) that joins the center of the patch to the nose tip and the head orientation .theta.=(.theta..sub.yaw, .theta..sub.pitch, .theta..sub.rol1). A number of such positive patches are extracted from each depth image. For negative patches however, there is no associated vector v and orientation .theta.. These extracted positive and negative patches are then used to train the Random Forests algorithm. Step 4003 determines (i.e. trains) the trained pose model by using a machine learning method according to the plurality of positive and negative patches and the ground truth rotations. In an example, the trained pose model is a forest structure comprising a plurality of binary tree structures, wherein each leaf of the binary tree structures of the forest structure is associated with values about rotation. The values about rotation may be determined according to at least one of the ground truth rotations. The machine learning method could be a random forest method (as described in Breiman, Leo. "Random forests." Machine learning 45.1 (2001): 5-32) for determining the forest structure.” [0098-0099] “In an embodiment of determining the trained pose model for determining a face pose, a set of patches (typically a few tens) are extracted from each training image (example patches are 5001-5006 shown in the FIG. 5). Patches that happen to lie on the face (face regions are marked in the training images) are considered `positive` patches and patches that do not lie on the face are `negative` patches. The ground truth poses of the face for each training image may be stored along with the patch information. The goal of the model is to then learn an association between the information in the patches and the expected output variable.” [0102]) Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the teachings of Narasimha into Goodrich, in order to have less constraints about displacement between any input images.
21. With reference to claim 17, Goodrich teaches the one or more machine learning models generate a segmentation vector that associates each pixel in the image with an indication of whether the pixel corresponds to a background or data representing the depiction of a person, the AR liquid element being applied further based on the segmentation vector. (“the image and depth data processing module 706 generates a segmentation mask based at least in part on the image data. In an embodiment, the image and depth data processing module 706 determines the segmentation mask using a convolutional neural network to perform dense prediction tasks where a prediction is made for every pixel to assign the pixel to a particular object class (e.g., face/portrait or background), and the segmentation mask is determined based on the groupings of the classified pixels (e.g., face/portrait or background). … the image and depth data processing module 706 performs background inpainting and blurring of the received image data using at least the segmentation mask to generate background inpainted image data. In an example, the image and depth data processing module 706 performs a background inpainting technique that eliminates the portrait (e.g., including the user's face) from the background and blurring the background to focus on the person in the frame.” [0161-0162] “a client device (e.g., the client device 102), receives a selection of a selectable graphical item from a plurality of selectable graphical items, the selectable graphical item corresponds an augmented reality content generator including a 3D effect. The client device captures image data using at least one camera of the client device. The client device generates depth data using a machine learning model based at least in part on the captured image data. The client device applies, to the image data and the depth data, the 3D effect based at least in part on the augmented reality content generator.” [0167])
22. Claim 18 is similar in scope to the claim 1, and thus is rejected under similar rationale. Goodrich additionally teaches A system comprising: at least one processor of a device; and a memory component having instructions stored thereon that, when executed by the at least one processor, cause the at least one processor to perform operations (“FIG. 1 is a block diagram showing an example messaging system 100 for exchanging data (e.g., messages and associated content) over a network. The messaging system 100 includes multiple instances of a client device 102, each of which hosts a number of applications including a messaging client application 104.” [0032] “The annotation system 206 is shown as including an image and depth data receiving module 702, a sensor data receiving module 704, an image and depth data processing module 706, a 3D effects module 708, a rendering module 710, a sharing module 712, and an augmented reality content generator module 714. The various modules of the annotation system 206 are configured to communicate with each other (e.g., via a bus, shared memory, or a switch). Any one or more of these modules may be implemented using one or more computer processors 750 (e.g., by configuring such one or more computer processors to perform functions described for that module) and hence may include one or more of the computer processors 750 (e.g., a set of processors provided by the client device 102).” [0130])
23. Claim 19 is similar in scope to the claim 1, and thus is rejected under similar rationale. Goodrich additionally teaches A non-transitory computer-readable storage medium having stored thereon instructions that, when executed by at least one processor of a device, cause the at least one processor to perform operations (“The method 800 may be embodied in computer-readable instructions for execution by one or more computer processors such that the operations of the method 800 may be performed in part or in whole by the messaging client application 104, particularly with respect to respective components of the annotation system 206 described above in FIG. 7; … the image and depth data receiving module 702 receives image data and depth data captured by an optical sensor (e.g., camera) of the client device 102.” [0142-0143])
24. Claim 20 is similar in scope to the claim 2, and thus is rejected under similar rationale.
Allowable Subject Matter
25. Claims 12-14 are objected to being dependent upon rejected base claims. The claims would be allowable if rewritten in independent form including all the limitations of the base claims and any intervening claims.
The following is a statement of reasons for the indication of allowable subject matter:
Regarding claim 12, the prior arts of record fails to either individually or in combination teach the claimed feature of: “as the dense depth reconstruction indicates that an individual pixel corresponding to a specified portion of the person has moved from a first position to a second position: determining that the specified portion is at a third distance that is less than the first distance; and modifying the AR liquid element in a first manner concurrently with applying a second visual effect to the portion of the image.”
Claims 13 and 14 are also objected to for depending from claim 12.
Conclusion
26. Any inquiry concerning this communication or earlier communications from the examiner should be directed to Michelle Chin whose telephone number is (571)270-3697. The examiner can normally be reached on Monday-Friday 8:00 AM-4:30 PM.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http:/Awww.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner's supervisor, Kent Chang can be reached on (571)272-7667. The fax phone number for the organization where this application or proceeding is assigned is (571)273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https:/Awww.uspto.gov/patents/apply/patent- center for more information about Patent Center and https:/Awww.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/MICHELLE CHIN/
Primary Examiner, Art Unit 2614