Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Response to Arguments
Applicant’s arguments, see pg. 8 of applicant’s arguments, filed 07/28/2026, with respect to the rejection(s) of claim(s) 1-5, 7-9, 11-20 under 103 have been fully considered and are persuasive. Therefore, the rejection has been withdrawn. However, upon further consideration, a new ground(s) of rejection is made in view of Carrier.
Allowable Subject Matter
Claims 5 and 16 are objected to as being dependent upon a rejected base claim, but would be allowable if rewritten in independent form including all of the limitations of the base claim and any intervening claims.
Claim Rejections - 35 USC § 102
(a)(1) the claimed invention was patented, described in a printed publication, or in public use, on sale, or otherwise available to the public before the effective filing date of the claimed invention.
(a)(2) the claimed invention was described in a patent issued under section 151, or in an application for patent published or deemed published under section 122(b), in which the patent or application, as the case may be, names another inventor and was effectively filed before the effective filing date of the claimed invention.
Claims 1-2, 8-9, 11, 13-14, and 17-18 are rejected under 35 U.S.C. 102 as being anticipated by Carrier (Pub No. US 20180125610 A1).
1. (Currently amended) A non-transitory computer readable medium comprising instructions that, when executed by a processing device (Carrier [0165]: “Various embodiments of the disclosure further disclose a non-transitory, computer-readable storage medium storing a set of instructions capable of being executed by a processor of a mobile telecommunications device, that, when executed by the processor”),
cause the processing device to perform operations comprising:
capturing a video comprising a plurality of frames of a face of an individual (Carrier figures 5-7 where it shows captured frames comprise a face of an individual. Carrier [0152]: “Also described herein is the use of continuous imaging (shooting) of the teeth. For example, rather than taking individual images, e.g., one at a time, the apparatus or method may be configured to general patient photos by using a continuous shooting mode. A rapid series of images may be taken while moving the mobile device. Movement can be guided by the apparatus, and may be from left to right, upper to lower, etc. From the user's perspective it may be similar to taking a video, but a series of images (still images) may be extracted by the apparatus.” The continuous shooting mode corresponds to the claimed “capturing a video comprising a plurality of frames”);
determining that the video fails to satisfy one or more quality criteria, wherein the one or more quality criteria comprise at least one of: a mouth shape criterion, a jaw pose criterion, or an amount of visible teeth criterion (Carrier [0154]: “As mentioned above, any of these variations may include detection of the teeth automatically, e.g., by machine learning. Detection of the patient teeth automatically may improve photo quality. In some variations, the machine learning (e.g., the machine learning framework provided by Apple with iOS 11) may be used to detect the presence of teeth when photos are taken, and to further guide the user. For example, the user may be alerted when the teeth are not visible, or to automatically select which predetermined view overly to use, to indicate if the angle is not correct, to indicate that the user is too close or too far from the patient, etc.” Figure 5B shows a mouth shape criterion by way of recommending the use of retractors. Figures 5C-5D shows jaw pose criterion and mouth space criterion by way of using on-screen overlays and guidance information. Carrier [0126]: “Any of these methods and apparatuses may further include reviewing the captured image on the mobile telecommunications device to confirm image quality and/or automatically accept/reject the image(s). For example, a method or apparatus may be configured to check the image quality of the captured image and displaying on the screen if the image quality is below a threshold for image quality”);
determining, based on which of the one or more quality criteria the video fails to satisfy, one or more actions to be performed by the individual, wherein the one or more actions comprise at least one of: change facial expression, smile, or open mouth; and providing guidance of the one or more actions to be performed by the individual to cause an updated video to satisfy the one or more quality criteria (Carrier [0154]: “As mentioned above, any of these variations may include detection of the teeth automatically, e.g., by machine learning. Detection of the patient teeth automatically may improve photo quality. In some variations, the machine learning (e.g., the machine learning framework provided by Apple with iOS 11) may be used to detect the presence of teeth when photos are taken, and to further guide the user. For example, the user may be alerted when the teeth are not visible, or to automatically select which predetermined view overly to use, to indicate if the angle is not correct, to indicate that the user is too close or too far from the patient, etc.” Carrier [0099]: “For example, the methods and apparatuses described herein may use a user's own handheld electronics apparatus having a camera (e.g., smartphone) and adapt it so that the user's device guides the user in taking high-quality images (e.g., at the correct aspect ratio/sizing, magnification, lighting, focus, etc.) of a predetermined sequence of orientations. In particular, these apparatuses and methods may include the use of an ‘overlay’ on a real-time image of the screen, providing immediate feedback on each of the desired orientations, which may also be used to adjust the lighting and/or focus, as described herein.
“An overlay may include an outline (e.g., a perspective view outline) of a set of teeth that may be used as a guide to assist in placement of the camera to capture the patient image. The overlay may be based on a generic image of teeth, or it may be customized to the user's teeth, or to a patient-specific category (by patient age and/or gender, and/or diagnosis, etc.). The overlay may be shown as partially transparent, or it may be solid, and/or shown in outline” Figure 5B provides an example guidance to have an open mouth by way of using retractors so that teeth become more visible. Figures 5C-5D provide an overlay as guidance for changing the facial expression or smiling, e.g. by way of using overlay 533 and the displayed message “Smile” in figure 5D.).
As per claim 17, this claim is similar in scope to limitations recited in claim 1, and thus is rejected under the same rationale.
As per claim 2, Carrier teaches the claimed:
2. The non-transitory computer readable medium of claim 1, the operations further comprising:
capturing the updated video comprising a second plurality of frames of the face of the individual after providing the guidance;
and determining that the updated video satisfies the one or more quality criteria. (Carrier [0152]: “Also described herein is the use of continuous imaging (shooting) of the teeth. For example, rather than taking individual images, e.g., one at a time, the apparatus or method may be configured to general patient photos by using a continuous shooting mode. A rapid series of images may be taken while moving the mobile device. Movement can be guided by the apparatus, and may be from left to right, upper to lower, etc. From the user's perspective it may be similar to taking a video, but a series of images (still images) may be extracted by the apparatus. For example, the apparatus may automatically review the images and match (or approximately match) views to the predetermined views, e.g., using the overlays as described above. The apparatus may select only those images having a sufficiently high quality. For example, blurry, dark or not optimally positioned photos may be automatically rejected. Similarly, multiple photos may be combined (by stitching, averaging, etc.” The different configurations and the selection of only high-quality photos comprise the updated video. The new configurations can be used to capture images that better fit the needs of the image analysis.).
As per claim 18, this claim is similar in scope to limitations recited in claim 2, and thus is rejected under the same rationale.
As per claim 8, Carrier teaches the claimed:
8. The non-transitory computer readable medium of claim 1, the operations further comprising: outputting a notice of which criteria of the one or more quality criteria are not satisfied and how to satisfy the one or more quality criteria. (Carrier fig. 6A shows a notification that the photo needs to be retaken because it does not meet a quality criterion. It describes a lack of focus, which is the specific criterion that needs to be corrected, and can be satisfied by refocusing the image.).
As per claim 9, Carrier teaches the claimed
9. The non-transitory computer readable medium of claim 1, wherein determining that the video fails to satisfy the one or more quality criteria and providing the guidance are performed during the capturing of the video. (Carrier [0032]: “Any of the methods and apparatuses described herein may guide a user in taking a series. The method or apparatus may provide audible and/or visual instructions to the user. In particular, as mentioned above, any of these apparatuses may include an overlay on the display (screen) of the mobile telecommunications device showing an outline that may be matched to guide the user in taking the image(s). The overlay may be shown as an outline in a solid and/or semi-transparent color. An overlay may be shown for each predetermined view. The user may observe the screen and, once the image shows the patient's anatomy approximately matching within the overlay, the image may be captured. Image capture may be manual (e.g., manually triggered for capture by the user activating a control, such as pushing a button to take the image) and/or automatic (e.g., detected by the system and automatically triggered to take the image when the overlay is matched with the corresponding patient anatomy). In general, capturing or triggering the capture of the image of the patient's teeth (and/or the patient's head) when the overlay approximately matches with the image of the patient's teeth in the display may refer to automatic capturing/automatic triggering, semi-automatic capturing/semi-automatic triggering, or manual capturing/manual triggering. Automatic triggering (e.g., automatic capturing) may refer to automatic capture of the image, e.g., taking one or more images when the patient's anatomy (e.g., teeth) show on the screen matches the overlay on the screen. Semi-automatic triggering (e.g., semi-automatic capturing) may refer to producing a signal, such as an audible sound and/or visual indicator (e.g., flashing, color change, etc.) when the patient's anatomy (e.g., teeth) shown on the screen matches the overlay on the screen. Manual triggering (e.g., manual capturing) may refer to the user manually taking the image, e.g., taking one or more images when the patient's anatomy (e.g., teeth) is shown on the screen to match the overlay.”).
As per claim 11, Carrier teaches the claimed:
11. The non-transitory computer readable medium of claim 1, the operations further comprising: determining facial landmarks of the face in one or more frames of the video;
determining at least one of a head position, a head orientation, a face angle, or a jaw position based on the facial landmarks;
and determining at least one of a) that the head position fails to satisfy a head position criterion, (Carrier [0150]: “Any of the apparatuses and methods described herein may also estimate between the patient and the camera of the mobile telecommunications device. For example, any of these methods and apparatuses may facial detection from the image to identify the patient's face; once identified, the size and position of the face (and/or any landmarks from the patient's face, such as eyes, nose, ears, lips, etc.) and may determine the approximate distance to the camera. This information may be used to guide the user in positioning the camera; e.g., instructing the user to get closer or further from the patient in order to take the images as described above. For example, in FIG. 15, the box 1505 on the screen identifies the automatically detected face/head of the patient; the face was identified using the face identification software available from the developer. FIG. 16 shows another example of face identification. In both examples, the method and apparatus may determine that the camera is too far away by comparing the size of the recognized face (the box region) to the size of the field. In both cases, the user may be provided instructions (audible, visual, text, etc.) to move the camera closer. In FIG. 15, the image to be taken is an anterior view of the teeth, and the user may be instructed to move much closer to focus on the teeth. In FIG. 16, the view may be a face (nonsmiling) image of the patient, and the user may be instructed to get closer. The distance to the patient may be estimated in any appropriate manner. For example, the distance may be derived from the size of the rectangle made around the patient's face or mouth, as mentioned (e.g., if it is too small, the camera is too far). The size of the ‘box’ identifying the face may be absolute (e.g., must be above a set value) or as a percent of the image field size.” Carrier claim 10: “10. The method of claim 1, wherein guiding the user comprises guiding the user to take the series of images comprises guiding the user to take at least one each of: an anterior view, a buccal view, an upper jaw view, and a lower jaw view.”).
b) that the head orientation fails to satisfy a head orientation criterion, c) that the face angle fails to satisfy a face angle criterion, or d) that the jaw position fails to satisfy a jaw position criterion.
As per claim 14, Carrier teaches the claimed:
14. The non-transitory computer readable medium of claim 1, the operations further comprising: determining an amount of visible teeth in the video; and determining whether the amount of visible teeth satisfies an amount of visible teeth criterion. (Carrier [0154]: “As mentioned above, any of these variations may include detection of the teeth automatically, e.g., by machine learning. Detection of the patient teeth automatically may improve photo quality. In some variations, the machine learning (e.g., the machine learning framework provided by Apple with iOS 11) may be used to detect the presence of teeth when photos are taken, and to further guide the user. For example, the user may be alerted when the teeth are not visible, or to automatically select which predetermined view overly to use, to indicate if the angle is not correct, to indicate that the user is too close or too far from the patient, etc.” The guidance to the user means the quality criteria of visible teeth count has been failed and needs to be met.).
As per claim 13 Carrier teaches the claimed:
13. The non-transitory computer readable medium of claim 1, the operations further comprising: detecting at least one of motion blur or camera focus associated with the video; and determining at least one of a) that the motion blur fails to satisfy a motion blur criterion or b) that the camera focus fails to satisfy a camera focus criterion. (Carrier [0167]: “In general, various embodiments of the disclosure further disclose a non-transitory, computer-readable storage medium storing a set of instructions capable of being executed by a processor of a mobile telecommunications device, that, when executed by the processor, causes the processor to display real-time images of the patient's teeth on a screen of the mobile telecommunications device and display an overlay comprising a cropping frame and an outline of teeth in one of an anterior view, a buccal view an upper jaw view, or a lower jaw view, wherein the overlay is displayed atop the images of the patient's teeth, and enable capturing of an image of the patient's teeth. The non-transitory, computer-readable storage medium, wherein the set of instructions, when executed by the processor, can further cause the processor to review the captured image and indicate on the screen if the captured image is out of focus and automatically crop the captured image as indicated by the cropping frame.”).
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claims 3-4, and 19-20 are rejected under 35 U.S.C. 103 as being unpatentable over Carrier as applied to claims above, and further in view of Ju (Pub No. US 20130315556 A1).
As per claim 3, Carrier alone fails to teach the claimed limitations.
However, Carrier in combination with Ju teaches the claimed:
3. The non-transitory computer readable medium of claim 2, the operations further comprising:
determining that one or more frames of the second plurality of frames of the updated video fail to satisfy the one or more quality criteria;
and removing the one or more frames from the updated video. (Ju teaches checking video frames for quality and dropping those that don’t meet it. Ju [0040]: “FIG. 4 is a diagram illustrating a third video recording example based on the proposed video recording apparatus 100 shown in FIG. 1. In this example, the input circuit 102 directly outputs the input video sequence V_IN as the first video sequence V_1 composed of video frames F.sub.1-F.sub.10, where the frame rate of the input video sequence V_IN is 120 Hz. As the video frames F.sub.2-F.sub.4 and F.sub.6-F.sub.9 include blurry image contents, the corresponding image quality metric values calculated by the image quality estimation circuit 104 would indicate that the video frames F.sub.2-F.sub.4 and F.sub.6-F.sub.9 have worse quality. Thus, the selection circuit 106 generates the second video sequence V_2 by selecting video frames F.sub.1, F.sub.5, F.sub.10 and dropping video frames F.sub.2-F.sub.4, F.sub.6-F.sub.9. In this example, the target frame rate of the output video sequence V_OUT is 30 Hz, which is lower than the frame rate of the input video sequence V_IN. Though the frame rate of the second video sequence V_2 is equal to the target frame rate of the output video sequence V_OUT due to the proposed image quality based video frame selection scheme, the interval between the image capture timing of the video frames F.sub.5 and F.sub.10 is not equal to an expected interval between image display timing of consecutive video frames (e.g., 1/30 second). …” Ju teaches dropping the frames of poor quality from the original video to make a second sequence.).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to use the removal of frames that don’t meet as taught by Ju with the system of Carrier in order to remove frames in a video sequence that don’t give sufficient information about the user according to the quality criteria from the images being analyzed.
As per claim 19, this claim is similar in scope to limitations recited in claim 3, and thus is rejected under the same rationale.
As per claim 4, Carrier alone fails to teaches the claimed limitations:
However, Carrier in combination with Ju teaches the claimed.
4. The non-transitory computer readable medium of claim 3, the operations further comprising: generating a replacement frame for at least one removed frame, wherein the replacement frame is generated based on a first frame preceding the removed frame and a second frame following the removed frame and comprises an intermediate state of the face between a first state of the face in the first frame and a second state of the face in the second frame. (Ju teaches interpolating a replacement for a removed frame based on at least one adjacent frame. These can include the preceding and following frames. Ju [0040]: “…and adding a new video frame F.sub.9' to the second video sequence V_2, where the interval between the image capture timing of the video frames F.sub.5 and F.sub.9' is equal to an expected interval between image display timing of consecutive video frames (e.g., 1/30 second). In one exemplary design, the video frame F.sub.9' is interpolated based on at least one of the adjacent selected video frames F.sub.5 and F.sub.10 in the second video sequence V_2. The video frame interpolation may adjust the weighting factors of the referenced video frames to obtain an interpolated video frame with good quality. …”).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to use the generations of replacement frames that fit with the quality criteria as taught by Ju with the system of Carrier in order to maintain the continuity of a video stream and use the examples of good-quality frames to interpolate new frames.
As per claim 20, this claim is similar in scope to limitations recited in claim 4, and thus is rejected under the same rationale.
Claim 7 is rejected under 35 U.S.C. 103 as being unpatentable over Carrier in view of Lee (US 20160092751 A1).
As per claim 7, Carrier alone does not explicitly teach the claimed limitations.
However, Carrier in combination with Zavesky teaches the claimed:
7. The non-transitory computer readable medium of claim 1, the operations further comprising: outputting a notice of the one or more quality criteria prior to beginning capturing of the video. (Zavesky [0038]: “Responsive to receiving the image data 110, the image processing system 102 may make a determination of a quality of a representation of one or more objects based on the image data 110. In particular, the image processing system 102 may identify a particular object based on the particular image capture setting and provide a notification as to the quality of the particular object. For example, when the particular image capture setting is the portrait setting, the image processing system 102 may identify a person (e.g., a face) based on the image data 110 and initiate a notification that indicates a quality of the representation of the person within the captured image. To illustrate, the notification may indicate that the image of the person is a high quality image.”
To measure the quality of the image being captured, the criteria for quality must be set beforehand and indicated so that a notification can be made if the standard is met.).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to use the indication of quality standards for the images as taught by Zavesky with the system of Carrier in order to show the user the criteria to be met and guide them on how to improve those metrics before the capturing a video so that the user keeps those criteria in mind in the process of filming.
Claim 12 is rejected under 35 U.S.C. 103 as being unpatentable over Carrier in view of Sachs (Pub No. US 20190122411 A1) and further in view of Wu (Pub No. US 20140371599 A1).
As per claim 12, Carrier alone does not explicitly teach the claimed limitations.
However, Carrier in combination with Sachs teaches the claimed:
12. The non-transitory computer readable medium of claim 1, the operations further comprising:
determining an optical flow between two or more frames of the video; (Sachs [0079]: “In accordance with some embodiments, a rig is generated for the static 3D model. The rig can be generated by applying a standard set FACS blend shapes to a mesh of the static 3D model of the head. The motion of one or more landmarks and/or 3D shapes in visible video can be tracked and the blend shapes of the static 3D model video recomputed based on the tracked landmarks and/or to provide a customized rig for the 3D model” The tracking, blending, and computation of the motion of the objects in the video is the optical flow.).
determining at least one of a head movement speed or a camera stability based on the optical flow; (The motion of the objects in Sachs includes the head, as describes above in [0141].).
Carrier and Sachs alone do not explicitly teach the claimed limitations.
However, Carrier in combination with Sachs and Wu teaches the claimed:
and determining at least one of a) that the camera stability fails to satisfy a camera stability criterion or b) that the head movement speed fails to satisfy a head movement speed criterion. (Wu [0056]: “As described herein, one or more processors of a computing device may be configured to identify patient behaviors from video information captured by camera 26. For example, the computing device may be configured to obtaining video information of patient motion captured over a period of time, such that the video information comprises a plurality of frames. The computing device may then receive, with respect to one or more frames of the plurality of frames, a selection of a sample area representative of an anatomical region (e.g., head 14, torso 16, arm 18A, or arm 18B). This sample area may be defined by user input and/or the one or more processors. The computing device may also analyze each of the other plurality of frames for respective areas corresponding to the sample area. The computing device can then calculate one or more movement parameters (e.g., velocity, angle of movement, or frequency of movement) of the anatomical region during the period of time from at least one difference between the sample area and one or more respective areas of at least a subset of the plurality of frames. The computing device may also be configured to compare the one or more movement parameters of the period of time to respective criteria for each of a plurality of predetermined patient behaviors (e.g., types of movements or movement disorders) and identify, based on the comparison, each one of the predetermined patient behaviors that occurred during the period of time.” Wu teaches comparing movement of a patient’s body parts. in an image to predetermined standards. These movements include velocity of a head area.).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to use the tracking of landmarks in a video as taught by Sachs with the system of Carrier in order to track the features of the head used to determine if an image meets a quality criterion.
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to use the tracking of the speed of the head in a video as taught by Wu with the system of Carrier modified by Sachs in order to track head speed as a criterion for images to determine if they meet a criterion of quality.
Claims 15 are rejected under 35 U.S.C. 103 as being unpatentable over Carrier in view of Xue (US 20180360567 A1).
As per claim 15, Carrer alone does not explicitly teach the claimed limitations.
However, Carrier in combination with Xue teaches the claimed:
15. The non-transitory computer readable medium of claim 1, wherein the updated video comprises a current condition of a dental site of the individual, the operations further comprising:
estimating a future condition of the dental site; and modifying the updated video by replacing the current condition of the dental site with the future condition of the dental site in the updated video. (Xue [0074]: “The apparatuses and/or methods (e.g., systems, devices, etc.) described below can be used with and/or integrated into an orthodontic treatment plan. The apparatuses and/or methods described herein may be used to segment a patient's teeth from a two-dimensional image and this segmentation information may be used to simulate, modify and/or choose between various orthodontic treatment plans. Segmenting the patient's teeth can be done automatically (e.g., using a computing device). For example, segmentation can be performed by a computing system automatically by evaluating data (such as three-dimensional scan, or a dental impression) of the patient's teeth or arch.” The treatment plan is the estimated future condition. Xue [0075]: “As described herein, an intraoral scanner may image a patient's dental arch and generate a virtual three-dimensional model of that dental arch. During an intraoral scan procedure (also referred to as a scan session), a user (e.g., a dental practitioner) of an intraoral scanner may generate multiple different images (also referred to as scans or medical images) of a dental site, model of a dental site, or other object. The images may be discrete images (e.g., point-and-shoot images) or frames from a video (e.g., a continuous scan). The three-dimensional scan can generate a 3D mesh of points representing the patient's arch, including the patient's teeth and gums. Further computer processing can segment or separate the 3D mesh of points into individual teeth and gums.” The future treatment plan can be applied to the updated video after the quality criteria have been applied.).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to use the modeling of an orthodontic treatment as taught by Xue with the system of Carrier in order to apply a simulated orthodontic treatment on a video after it has been modified to meet quality criteria.
Conclusion
Any inquiry concerning this communication or earlier communications from the examiner should be directed to THOMAS JOHN FOSTER whose telephone number is (571)272-5053. The examiner can normally be reached Mon, Fri 8:30-6. Tues-Thurs 7:30-5.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Daniel Hajnik can be reached at 571-272-7642. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/THOMAS JOHN FOSTER/ Examiner, Art Unit 2616
/DANIEL F HAJNIK/ Supervisory Patent Examiner, Art Unit 2616