DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Response to Preliminary Amendment
This Office Action is responsive to communications filed on 11/25/2024 and a preliminary amendment also filed on 11/25/2024. Applicant has amended the specification to amend the priority paragraph. Applicant has also canceled original claims 1-20 without prejudice or disclaimer of the subject matter thereof, and has added new claims 21-40, of which claims 21 and 31 are independent. Applicant submits no new matter has been added. Claims 21-40 are pending in the instant application. An Office Action on the merits follows here below.
Priority
This application discloses and claims only subject matter disclosed in prior application number 18/644,261, filed 04/24/2024, and names the inventor or at least one joint inventor named in the prior application. Accordingly, this application has been examined as a continuation.
Information Disclosure Statement
The information disclosure statement (IDS) submitted on 04/09/2025 is in compliance with the provisions of 37 CFR 1.97. Accordingly, the information disclosure statement is being considered by the examiner.
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows:
1. Determining the scope and contents of the prior art.
2. Ascertaining the differences between the prior art and the claims at issue.
3. Resolving the level of ordinary skill in the pertinent art.
4. Considering objective evidence present in the application indicating obviousness or nonobviousness.
Claims 21, 22, 23, 24, 26-28, 31, 37, 38 and 40 are rejected under 35 U.S.C. 103 as being unpatentable over Zhang (US 20210027485 A1) in combination with Kadam “FVEstimator” FVEstimator: A novel food volume estimator Wellness model for calorie measurement and healthy living, Measurement, Volume 198, July 2022, 111294, ISSN 0263-2241, https://doi.org/10.1016/j.measurement.2022.111294.
Regarding Claim 21 (New): Zhang discloses a non-transitory computer-readable medium including instructions that, when executed by at least one processor, cause the at least one processor to perform operations for providing a personalized dietary recommendation, the operations comprising: (Refer to para [059]; “Other embodiments of these and other aspects of the disclosure include corresponding systems, apparatus, and computer programs, configured to perform the actions of the methods, encoded on computer storage devices. A system of one or more computers can be so configured by virtue of software, firmware, hardware, or a combination of them installed on the system that in operation cause the system to perform the actions. One or more computer programs can be so configured by virtue having instructions that, when executed by data processing apparatus, cause the apparatus to perform the actions.”) capturing an RGB image using an integrated camera (Refer to para [129]; “The process includes obtaining image data (step 502), such as from one or more cameras or other sensors located to capture data about a monitored area.”) inputting the RGB image into an instance detection network configured to detect food items (Refer to para [025 and 026]; “… the one or more machine learning models comprise a convolutional neural network. In some implementations, the one or more machine learning models comprise a neural network including a region proposal network portion configured to identify regions within an image and an object detection network portion configured to classify the regions identified by the region proposal network portion.”) segmenting a plurality of food items from the RGB image into a plurality of masks, the plurality of masks representing individual food items (Refer to para [123]; “This can be done by, for example, running an image segmentation process on the image 300 to identify high-contrast or high-sharpness boundaries in the image 300. The object detection model can indicate predicted regions covering a majority of the objects or a center of the objects. From the center, the shading or other marking may extend outward to the boundaries noted in the segmentation process to cover visible surfaces of an object.”) classifying a particular food item among the individual food items using a multimodal large language model (Refer to para [013]; “the one or more machine learning models have been trained to detect a plurality of different types of objects and to indicate a status of at least one of the types of objects; and the output of the one or more machine learning models indicates (i) locations of identified objects in the image data representing the image of the monitored area, (ii) an object status classification for at least one of the identified objects, and (iii) confidence scores for the identification of the objects and/or the object status classifications.”).
Zhang does not expressly disclose estimating a volume of a particular food item.
Kadam teaches food volume estimation using instance segmentation using mask-based RCNN.
Further, Kadam teaches “capturing the accurate food volume consumed inside or outside the house. This method will assist patients/sports persons/weight watchers in measuring the food volume and calculating the calories.” … such that the image processor is capable to calculate estimating a volume of the particular food item by overlaying an RGB image with a depth-map (Refer to page 2, right column, Section 3, para [003] “have used SSD-Mobile-net object detection model to find Volume and Density of food and further to calculate the weight of Food item, weight to a calorie is converted using a standard chart. The object recognition accuracy was 0.99, R2 = 0.95, RMSE = 43, MAE = 32 and average error rate was 9%. The dataset consisted of a total of 633 original images, 11 categories, averaging nearly 60 images in each category, and a COCO dataset. Other study [27] uses CNN model with 3D reconstruction while Mean absolute error (Classification) = 9%, Food volume average accuracy = 92%. The dataset used was Food X, RGB-D images, and an extensive 3D food model dataset. The authors in [28] used UNet, VNet, which gave volume testing accuracy as 92.29%. They used eight food categories: banana, apple, burger, cake, pizza, orange, rice, and donuts.”) to create a point cloud, wherein the depth map is created by monocular depth estimation (Refer to page 3, right column, Section 6, para [ 002]; “An image from this dataset is now segmented by instance segmentation MaskRCNN. Two output images are created after applying the MaskRCNN algorithm — the bounding box is marked across the food item, and masks of the food item are generated.”) and generating the personalized dietary recommendation based on the classified food item, the estimated volume, and the user profile data (Refer to page 11, left column, Section 9, para [001]; “The proposed FVEstimator Wellness model finds the dimensions — height and diameter of the bowl, which is equivalent to the dimensions of the food contained in it. The work presents the relevance of instance segmentation with Mask RCNN in food image segmentation. Works [20] have already proven the results of Mask RCNN in food image segmentation. Our work establishes its accuracy on the dataset with regular and amorphous shapes. The model is further enhanced with a pre-trained RESNET network due to transfer learning. From the dimensions obtained, the volume of the food is calculated with an accuracy of 90.46 %. Similarly, food items that take the shape of the bowl and keep the shape are given a convex shape. The convex shaped food items dimensions are calculated similarly with an accuracy of 90.0%..”).
Before the effective filing date of the claimed invention, it would have been obvious to a person of ordinary skill in the art to modify Zhang by adding an image processor for “images-based automated food volume estimation using deep learning methods…” as taught above by Kadam.
The suggestion/motivation for combining the teachings of Zhang and Kadam would have been in order to “enhance automatic image-based dietary monitoring systems, which will help the patients, athletes, weight-watchers, etc., to record their diet and maintain a healthy lifestyle by tracking the calorie count by correctly identifying and estimating the portion size of the food they are eating.” (Section 11. Conclusion, para [001], Kadam).
Therefore, it would have been obvious to one of ordinary skill in the art to combine the teachings of Zhang and Kadam in order to obtain the specified claimed elements of Claim 21. It is for at least the aforementioned reasons that the Examiner has reached a conclusion of obviousness with respect to the claim in question.
Regarding Claim 22: (New) Zhang discloses wherein the operations further comprise using the multimodal large language model to analyze the user profile data (Refer to para [115]; “In some implementations, video feeds or sequences of images from the cameras 110a, 110b can be analyzed by the models 123 to determine if the types of motion or types of changes in the monitored area represent conditions that need attention. This can be helpful to detect events or movement patterns that are unusual for the monitored area.”). The Examiner maintains that analysis performed by Zhang more than fairly utilizes a large language model as detailed at para [122]; “neural network also indicates a confidence score indicating a likelihood that the model assigns for the prediction being correct.” which one of ordinary skill in the art understands that the neural network is capable to operate analysis of data.
Regarding Claim 23: (New) Kadam teaches wherein the user profile data comprises personal data of a user and dietary log entries (Refer to page 10, Section 9.3; para [001]; “The study of the current dietary management system reveals the significance of automated image-based dietary monitoring systems over manual logging methods [22]. Food volume estimation in dietary management persists as a challenge for food items of different shapes, regular, irregular, solid, and amorphous. The presented work carried out a survey of deep learning techniques [15,19,20] in the food volume estimation with complete description to its datasets, food types, and performance measures.”).
Regarding Claim 24: (New) Kadam teaches wherein the operations further comprise estimating a calorie amount of the particular food item using the estimated volume and a nutritional database (Refer to page 3, Section 6, para [001]; “The work in this paper focuses on automatic dietary monitoring by capturing food images, estimating its volumetric weight, and then estimating the total calories of the entire meal partaken by the individual. Fig. 1 represents the obesity determinants.”).
Regarding Claim 26 (New): Kadam teaches wherein the personalized dietary recommendation includes a recommendation related to a medical diet (Refer to Section 2; “This problem came into highlight by discussing with a dietician who recommends specific diets to her patients using terms like — one bowl of food or 300 gms of the food item. Now one bowl is very subjective as bowl size at one patient’s home will be different from the other. So, when the dietician recommends the diet, a tool that estimates portion size by capturing the image of the meal will help the patient verify that the meal portion is according to the suggested diet.”).
Regarding Claim 27 (New): Kadam teaches wherein the personalized dietary recommendation includes a recommended portion size of the particular food item (Refer to page 11, Section 11; “The work opens the research opportunities in images-based automated food volume estimation using deep learning methods. This will help enhance automatic image-based dietary monitoring systems, which will help the patients, athletes, weight-watchers, etc., to record their diet and maintain a healthy lifestyle by tracking the calorie count by correctly identifying and estimating the portion size of the food they are eating. The cuisine and food around the world consist of a variety of food items with a variety of shapes. Our first contribution is the creation of the dataset of a variety of food shapes, regular and irregular alike. The dataset also contains tableware that defines the amorphous food shape. The Food volume estimation is done using the volume estimation of tableware containing the food, and validation is done by actually measuring the dimensions of the tableware containing the food.”).
Regarding Claim 28: (New) Kadam teaches wherein the personalized dietary recommendation includes a recommended related to the nutritional information of the particular food item (Refer to page 11, Section 11; “The work opens the research opportunities in images-based automated food volume estimation using deep learning methods. This will help enhance automatic image-based dietary monitoring systems, which will help the patients, athletes, weight-watchers, etc., to record their diet and maintain a healthy lifestyle by tracking the calorie count by correctly identifying and estimating the portion size of the food they are eating. The cuisine and food around the world consist of a variety of food items with a variety of shapes. Our first contribution is the creation of the dataset of a variety of food shapes, regular and irregular alike. The dataset also contains tableware that defines the amorphous food shape. The Food volume estimation is done using the volume estimation of tableware containing the food, and validation is done by actually measuring the dimensions of the tableware containing the food.”).
Regarding Claim 31 (New): Zhang discloses a computer-implemented method for providing a personalized dietary recommendation, the operations comprising: (Refer to para [090]; “This determination of status can be made for any and/or all of the types of objects that the models 123 are trained to detect. For example, for a person, the models 123 may be trained to classify the activity of the person, e.g., to provide outputs indicating likelihoods whether the person is eating, waiting, ordering, passing through, etc.”) capturing an RGB image using an integrated camera (Refer to para [092] and [098]); “As noted above, models 123 can be trained to detect the presence of an object in an image, to determine the location of the object in the image, and to determine a status classification for the object from among multiple different possible status classifications.. specific cameras 110a, 110b.”)
inputting the RGB image into an instance detection network configured to detect food items (Refer to para [025 and 026]; “… the one or more machine learning models comprise a convolutional neural network. In some implementations, the one or more machine learning models comprise a neural network including a region proposal network portion configured to identify regions within an image and an object detection network portion configured to classify the regions identified by the region proposal network portion.”) segmenting a plurality of food items from the RGB image into a plurality of masks, the plurality of masks representing individual food items (Refer to para [123]; “This can be done by, for example, running an image segmentation process on the image 300 to identify high-contrast or high-sharpness boundaries in the image 300. The object detection model can indicate predicted regions covering a majority of the objects or a center of the objects. From the center, the shading or other marking may extend outward to the boundaries noted in the segmentation process to cover visible surfaces of an object.”) classifying a particular food item among the individual food items using a multimodal large language model (Refer to para [013]; “the one or more machine learning models have been trained to detect a plurality of different types of objects and to indicate a status of at least one of the types of objects; and the output of the one or more machine learning models indicates (i) locations of identified objects in the image data representing the image of the monitored area, (ii) an object status classification for at least one of the identified objects, and (iii) confidence scores for the identification of the objects and/or the object status classifications.”).
Zhang does not expressly disclose generating a personalized dietary recommendation based on the classified food item.”
Kadam teaches food volume estimation using instance segmentation using mask-based RCNN.
Further, Kadam teaches “capturing the accurate food volume consumed inside or outside the house. This method will assist patients/sports persons/weight watchers in measuring the food volume and calculating the calories.” … such that the image processor is capable to calculate estimating a volume of the particular food item by overlaying an RGB image associated with a depth-map (Refer to page 2, right column, Section 3, para [003] “have used SSD-Mobile-net object detection model to find Volume and Density of food and further to calculate the weight of Food item, weight to a calorie is converted using a standard chart. The object recognition accuracy was 0.99, R2 = 0.95, RMSE = 43, MAE = 32 and average error rate was 9%. The dataset consisted of a total of 633 original images, 11 categories, averaging nearly 60 images in each category, and a COCO dataset. Other study [27] uses CNN model with 3D reconstruction while Mean absolute error (Classification) = 9%, Food volume average accuracy = 92%. The dataset used was Food X, RGB-D images, and an extensive 3D food model dataset. The authors in [28] used UNet, VNet, which gave volume testing accuracy as 92.29%. They used eight food categories: banana, apple, burger, cake, pizza, orange, rice, and donuts.”) to create a point cloud (Refer to page 3, right column, Section 6, para [ 002]; “An image from this dataset is now segmented by instance segmentation MaskRCNN. Two output images are created after applying the MaskRCNN algorithm — the bounding box is marked across the food item, and masks of the food item are generated.”) wherein the depth map is created by monocular depth estimation; (Refer to page 2, left column, para [001]; “Based on the computation of images, the methods for FVE are — Stereo-based approach, Model-based approach, Perspective transformation approach, Depth camera-based approach, and Deep learning approach. An in depth comparison of the Model-based approach (3D reconstruction) and Deep Learning was made by [14]. The % error range for Deep learning approach is : 1.62 to 9.28 [15], and for 3D reconstruction : 0.83 to 12.3. This % error is for canned and uncooked food items.”) and generating the personalized dietary recommendation based on the classified food item, the estimated volume and user profile data (Refer to page 11, left column, Section 9, para [001]; “The proposed FVEstimator Wellness model finds the dimensions — height and diameter of the bowl, which is equivalent to the dimensions of the food contained in it. The work presents the relevance of instance segmentation with Mask RCNN in food image segmentation. Works [20] have already proven the results of Mask RCNN in food image segmentation. Our work establishes its accuracy on the dataset with regular and amorphous shapes. The model is further enhanced with a pre-trained RESNET network due to transfer learning. From the dimensions obtained, the volume of the food is calculated with an accuracy of 90.46 %. Similarly, food items that take the shape of the bowl and keep the shape are given a convex shape. The convex shaped food items dimensions are calculated similarly with an accuracy of 90.0%..”).
Before the effective filing date of the claimed invention, it would have been obvious to a person of ordinary skill in the art to modify Zhang by adding an image processor for “images-based automated food volume estimation using deep learning methods…” as taught above by Kadam.
The suggestion/motivation for combining the teachings of Zhang and Kadam would have been in order to “… enhance automatic image-based dietary monitoring systems, which will help the patients, athletes, weight-watchers, etc., to record their diet and maintain a healthy lifestyle by tracking the calorie count by correctly identifying and estimating the portion size of the food they are eating.” (Section 11. Conclusion, para [001], Kadam).
Therefore, it would have been obvious to one of ordinary skill in the art to combine the teachings of Zhang and Kadam in order to obtain the specified claimed elements of Claim 31. It is for at least the aforementioned reasons that the Examiner has reached a conclusion of obviousness with respect to the claim in question.
Regarding Claim 37: (New) Zhang discloses wherein the instance detection network is trained on a plurality of reference food items using a neural network (Refer to para [115]; “In some implementations, video feeds or sequences of images from the cameras 110a, 110b can be analyzed by the models 123 to determine if the types of motion or types of changes in the monitored area represent conditions that need attention. This can be helpful to detect events or movement patterns that are unusual for the monitored area.”). The Examiner maintains that analysis performed by Zhang more than fairly utilizes a large language model as detailed at para [122]; “neural network also indicates a confidence score indicating a likelihood that the model assigns for the prediction being correct.” which one of ordinary skill in the art understands that the neural network is capable to operate analysis of data.
Regarding Claim 38: (New) Kadam teaches wherein the operations further comprise estimating a calorie amount of the particular food item using the estimated volume and a nutritional database (Refer to page 3, Section 6, para [001]; “The work in this paper focuses on automatic dietary monitoring by capturing food images, estimating its volumetric weight, and then estimating the total calories of the entire meal partaken by the individual. Fig. 1 represents the obesity determinants.”).
Regarding Claim 40: (New) Kadam teaches wherein the personalized dietary recommendation includes a recommendation related to nutritional information of the particular food item (Refer to page 11, Section 11; “The work opens the research opportunities in images-based automated food volume estimation using deep learning methods. This will help enhance automatic image-based dietary monitoring systems, which will help the patients, athletes, weight-watchers, etc., to record their diet and maintain a healthy lifestyle by tracking the calorie count by correctly identifying and estimating the portion size of the food they are eating. The cuisine and food around the world consist of a variety of food items with a variety of shapes. Our first contribution is the creation of the dataset of a variety of food shapes, regular and irregular alike. The dataset also contains tableware that defines the amorphous food shape. The Food volume estimation is done using the volume estimation of tableware containing the food, and validation is done by actually measuring the dimensions of the tableware containing the food.”).
Allowable Subject Matter
Claims 25, 29, 30, 32-36 and 39 are objected to as being dependent upon a rejected base claim, but would be allowable if rewritten in independent form including all of the limitations of the base claim and any intervening claims.
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure:
Adachi “DepthGrillCAM: A Mobile Application for Real-Time Eating Action Recording Using RGB-D Images” discloses “In this study, we proposed a mobile food recognition system that can check calories during a meal in real time by combining calorie estimation and eating action recognition with depth information that can be obtained by a depth sensor mounted on the front of a smartphone used for face recognition. The experiments show that in the situation of eating grilled meat, the proposed method can improve the accuracy of calorie estimation by up to 28% compared to the conventional method, GrillCam [1], and can recognize the correct meal category with 6.67 times higher accuracy in eating action recognition. In the future, we will expand the dataset and aim to improve the accuracy of calorie estimation by considering the three-dimensional shape of the food.”
Muhsin (US 20230222887 A1) discloses “one or more infant safety models can be applied. In response to detecting an infant, the camera system can apply one or more infant safety models that outputs a model result. The camera system can invoke (which can be invoked on a hardware accelerator) an infant position model based on the captured data. The infant position model can output a classification result. In some aspects, the infant position model can be or include a CNN. In response to detecting an infant, the camera system can invoke a facial feature extraction model based on second image data where the facial feature extraction model outputs a facial feature vector. The camera system can execute a query of a facial features database based on the facial feature vector where executing the query indicates that the facial feature vector is not present in the facial features database. An infant safety model can be an infant color detection model. In some aspects, the model result can include coordinates of a boundary region identifying an infant object in the image data. As described herein, the camera system can invoke a loud noise detection model based on the audio data where the loud noise detection model can output a classification result. Some of the aspects described herein can include any of the following features, which can be applied in different settings. In some aspects, a camera system can have local storage for an image and/or video feed. In some aspects, remote access of the local storage may be restricted and/or limited. In some aspects, the camera system can use a calibration factor which can be useful for correcting color drift in the image data from a camera. In some aspects, the camera system can add or remove filters on camera to provide certain effects. The camera system may include infrared filters. In some aspects, the monitoring system can monitor food intake of subject and/or estimate calories. In some aspects, the monitoring system can detect mask wearing (such as wearing or not wearing an oxygen mask).”
Any inquiry concerning this communication or earlier communications from the examiner should be directed to MIA M THOMAS whose telephone number is (571)270-1583. The examiner can normally be reached M-Th 8:30am-4:30pm.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Edward (Ed) Urban can be reached on 572-272-7899. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
MIA M. THOMAS
Primary Examiner
Art Unit 2665
/MIA M THOMAS/Primary Examiner
Art Unit 2665