DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Status of the Claims
Claims 1-20 are pending for examination.
Claims 1, 19 and 20 are independent Claims.
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claim(s) ) 1-2, 12 and 14-20 is/are rejected under 35 U.S.C. 102(a)(1) as being anticipated by Zueck (U.S. 8,515,497 hereinafter Zueck) in view of Wright (U.S. 2020/0394012 hereinafter Wright).
As Claim 1, Zueck teaches one or more non-transitory computer-readable media comprising computer-executable instructions that, when executed by one or more processors, cause performance of operations, comprising:
detecting a storage, of an image as an image file, triggered by a capturing of the image by an image-capture function of a mobile computing device (Zueck (col. 6 line 32-35, fig. 6), “at block 202 where at least one image is captured with the camera 26. The method then proceeds to block 204 where the images are stored in the memory 122”);
and contemporaneously with detecting the storage of the image file, activating an audio recording function to capture an audio recording; (Zueck (col. 6 line 36-47 and 54-62, fig. 6 item 208), “At block 206, the switch 22 is activated that triggers a voice generated picture file mode after the user activates the switch. In one illustrative embodiment the switch automatically activates various steps associated with the voice generated picture file mode”, “In one embodiment, the switch 22 operations at block 206 are bypassed and the method proceeds to block 208” and “Typically the processing of the voice message occurs when an image or video is captured by the camera 26”);
receiving the audio recording captured by the audio recording function (Zueck (col. 6 line 36-40, fig. 6), “At block 206, the switch22 is activated that triggers a voice generated picture file mode after the user activates the switch. In one illustrative embodiment the switch automatically activates various steps associated with the voice generated picture file mode”); and
Zueck may not explicitly disclose:
detecting one or more features of the image;
based on the one or more features of the image,
storing (a) the audio recording as an audio file and (b) an association between the audio file and the image file.
Wright teaches:
detecting one or more features of the image (Wright (¶0223 middle portion), “the system may key annotations and audio off of actual recognition of objects meaning once an object or feature point is recognized in the video, an audio or annotation sequence may be initiated”);
based on the one or more features of the image (Wright (¶0223 middle portion), “the system may key annotations and audio off of actual recognition of objects meaning once an object or feature point is recognized in the video, an audio or annotation sequence may be initiated”),
storing (a) the audio recording as an audio file (Wright (¶0272 middle portion), “detecting a first object within the live video stream; and anchoring the audio-visual communication to the detected first object,”) and (b) an association between the audio file and the image file (Wright (¶0272 middle portion), “storing an indication of the anchoring of the audio-visual communication to the detected first object independent of the live video stream.”).
Zueck discloses a system/method to capture image and audio and associate the transcript of the audio with the image. Wright discloses a system to attach captured audio file with the image file. It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify audio transcript of Zueck in view of Gafni instead be an audio file taught by Wright, with a reasonable expectation of success. The motivation would be to conveniently allow “once an object or feature point is recognized in the video, an audio or annotation sequence may be initiated.” (Arnold (¶0223 middle portion)).
As Claim 2, besides claim 1, Zueck in view of Wright teaches wherein the operations further comprise:
responsive to detecting the storage of the image, activating a voice command function (Zueck (col. 6 line 36-40, fig. 6), “At block 206, the switch22 is activated that triggers a voice generated picture file mode after the user activates the switch. In one illustrative embodiment the switch automatically activates various steps associated with the voice generated picture file mode”);
requesting, via the voice command function, an input from a user (Zueck (col. 6 line 36-37), “At block 206, the switch22 is activated that triggers a voice generated picture file mode after the user activates the switch”);
capturing the audio recording via the audio recording function, wherein the audio recording comprises the input from the user (Zueck (col 6 line 60-63), “Typically the processing of the voice message occurs when an image or video is captured by the camera 26. At block 210, the text-based file identifier is associated with the image that was captured by the camera 26”).
As Claim 12, besides Claim 1, Zueck in view of Wright teaches wherein the operations further comprise:
responsive at least in part to detecting the storage of the image, outputting a request comprising a description of a set of requested content for the audio recording to be captured via the audio recording function, wherein the description comprises at least of:
a text description of the set of requested content, or
an auditory description of the set of requested content (Zueck (col 6 line 60-63), “voice message”);
capturing the audio recording via the audio recording function, wherein the audio recording comprises a voice input from a user (Zueck (col 6 line 60-63), “Typically the processing of the voice message occurs when an image or video is captured by the camera 26. At block 210, the text-based file identifier is associated with the image that was captured by the camera 26”).
As Claim 14, besides Claim 1, Zueck in view of Wright teaches wherein the operations further comprise: responsive to detecting the storage of the image:
activating an image annotation function for associating an annotation with the image (Zueck (col. 6 line 36-40, fig. 6), “At block 206, the switch22 is activated that triggers a voice generated picture file mode after the user activates the switch. In one illustrative embodiment the switch automatically activates various steps associated with the voice generated picture file mode”);
generating the annotation (Zueck (col 6 line 60-63), “Typically the processing of the voice message occurs when an image or video is captured by the camera 26. At block 210, the text-based file identifier is associated with the image that was captured by the camera 26”); and
storing the annotation in association with the image (Zueck (col 6 line 60-63), “Typically the processing of the voice message occurs when an image or video is captured by the camera 26. At block 210, the text-based file identifier is associated with the image that was captured by the camera 26”).
As Claim 15, besides Claim 14, Zueck in view of Wright teaches wherein generating the annotation comprises:
applying the annotation to the image (Zueck (col 6 line 60-63), “Typically the processing of the voice message occurs when an image or video is captured by the camera 26. At block 210, the text-based file identifier is associated with the image that was captured by the camera 26”).
As Claim 16, besides Claim 14, Zueck in view of Wright teaches wherein generating the annotation comprises:
generating the annotation based on the image, wherein the annotation comprises a description of the image (Zueck (col 6 line 60-63), “Typically the processing of the voice message occurs when an image or video is captured by the camera 26. At block 210, the text-based file identifier is associated with the image that was captured by the camera 26”).
As Claim 17, besides Claim 14, Zueck in view of Wright teaches wherein the operations further comprise: capturing the audio recording via the audio recording function (Zueck (col 6 line 60-63), “voice message”); and generating the annotation based on the audio recording (Zueck (col 6 line 60-63), “Typically the processing of the voice message occurs when an image or video is captured by the camera 26. At block 210, the text-based file identifier is associated with the image that was captured by the camera 26”).
As Claim 18, besides Claim 17, Zueck in view of Wright teaches wherein generating the annotation based on the audio recording comprises: generating a text annotation based on the audio recording (Zueck (col 6 line 60-63), “Typically the processing of the voice message occurs when an image or video is captured by the camera 26. At block 210, the text-based file identifier is associated with the image that was captured by the camera 26”).
As Claim 19 and 20, the Claims are rejected for the same reasons as Claim 1.
Claim(s) 3-6, 8-9, 11 and 13 is/are rejected under 35 U.S.C. 103 as being unpatentable over Zueck in view of Wright in further view of Gafni et al. (U.S. 2021/0019573 hereinafter Gafni).
As Claim 3, besides Claim 1, Zueck in view of Wright teaches:
activating the audio recording function to capture the audio recording (Zueck (col. 6 line 36-40, fig. 6), “At block 206, the switch22 is activated that triggers a voice generated picture file mode after the user activates the switch. In one illustrative embodiment the switch automatically activates various steps associated with the voice generated picture file mode”).
Zueck may not explicitly disclose:
wherein activating the audio recording function to capture the audio recording responsive at least in part to detecting the storage of the image comprises:
responsive to detecting the storage of the image, executing a machine learning model, wherein executing the machine learning model comprises:
determining that the image meets a similarity threshold with a training dataset comprising a set of training images, and
responsive to determining that the image meets the similarity threshold with the training dataset,
Gafni teaches:
wherein activating the audio recording function to capture the audio recording responsive at least in part to detecting the storage of the image comprises:
responsive to detecting the storage of the image, executing a machine learning model, wherein executing the machine learning model comprises:
determining that the image meets a similarity threshold (Gafni (¶0081 last 6), “the electronic device may identify and select at least a subset of the objects (operation 214) based at least in part on analysis of the image. Next, the electronic device may receive classifications (operation 216) for objects in the subset, such as names for the objects or information that specifies types of objects”) with a training dataset comprising a set of training images (Gafni (¶0114 line 5-10), “multiple images may be used in one or more ways. For example, multiple images of a scene may be acquired from different points of view in order to train object detection and classification”), and
responsive to determining that the image meets the similarity threshold with the training dataset (Gafni (¶0081 last 6), “the electronic device may identify and select at least a subset of the objects (operation 214) based at least in part on analysis of the image. Next, the electronic device may receive classifications (operation 216) for objects in the subset, such as names for the objects or information that specifies types of objects”) (Gafni (¶0107 line 1-6), “electronic device 110-1 may ask the user for feedback on the determined one or more inspection criteria (which are, in essence, recommended inspection criteria). For example, the user may be asked to confirm or accept a recommended inspection criterion, or to modify or change it”),
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify image of Zueck instead be an image analysis module taught by Gafni, with a reasonable expectation of success. The motivation would be to “allow a
user to interactively specify the information needed to simply and efficiently generate and distribute an application that leverage image acquisition capabilities of the second electronic devices, image analysis and logical rules (the one or more inspection criteria)” (Gafni (¶0049 line 1-6)).
As Claim 4, besides Claim 3, Zueck in view of Wright in further view of Gafni teaches wherein the set of training images are associated with a set of descriptive information, wherein executing the machine learning model further comprises:
responsive to determining that the image meets the similarity threshold with the training dataset (Gafni (¶0081 last 6), “the electronic device may identify and select at least a subset of the objects (operation 214) based at least in part on analysis of the image. Next, the electronic device may receive classifications (operation 216) for objects in the subset, such as names for the objects or information that specifies types of objects”), associating the set of descriptive information with the image (Gafni (¶0086 line 1-5), “the electronic device may provide recommended classifications for the objects in the subset, where the received classifications for the objects in the subset are based at least in part on the recommended classifications”)
As Claim 5, besides Claim 4, Zueck in view of Wright in further view of Gafni teaches wherein executing the machine learning model further comprises:
providing the set of descriptive information as an audible or visual output of the mobile computing device (Gafni (¶0103 line 9-11, fig. 11 item 1110), “Note that electronic device 110-1 may present recommended classifications 912 and 1110, which are automatically determined by electronic 110-1 (and/or computer 120)”).
As Claim 6, besides Claim 4, Zueck in view of Wright in further view of Gafni teaches wherein executing the machine learning model further comprises: annotating the image with the set of descriptive information (Gafni (¶0103 line 9-11, fig. 11 item 1110), “Note that electronic device 110-1 may present recommended classifications 912 and 1110, which are automatically determined by electronic 110-1 (and/or computer 120)”).
As Claim 8, besides Claim 3, Zueck in view of Gafni teaches wherein the operations further comprise:
training the machine learning model to activate the audio recording function (Gafni (¶0107 line 1-6), “electronic device 110-1 may ask the user for feedback on the determined one or more inspection criteria (which are, in essence, recommended inspection criteria). For example, the user may be asked to confirm or accept a recommended inspection criterion, or to modify or change it”) when images meet the similarity threshold with the training dataset (Gafni (¶0081 last 6), “the electronic device may identify and select at least a subset of the objects (operation 214) based at least in part on analysis of the image. Next, the electronic device may receive classifications (operation 216) for objects in the subset, such as names for the objects or information that specifies types of objects”).
As Claim 9, besides Claim 1, Zueck in view of Wright teaches:
activating the audio recording function to capture the audio recording (Zueck (col. 6 line 36-40, fig. 6), “At block 206, the switch22 is activated that triggers a voice generated picture file mode after the user activates the switch. In one illustrative embodiment the switch automatically activates various steps associated with the voice generated picture file mode”).
Zueck in view of Wright may not explicitly disclose:
wherein detecting one or more features of the image comprises:
executing a machine learning model, wherein executing the machine learning model comprises:
determining, based on a comparison of the image to a training dataset comprising a set of training images, that the image comprises a particular feature, and responsive to determining that the image comprises the particular feature.
Gafni teaches:
wherein detecting one or more features of the image comprises:
executing a machine learning model, wherein executing the machine learning model comprises:
determining, based on a comparison of the image to a training dataset comprising a set of training images, that the image comprises a particular feature, and responsive to determining that the image comprises the particular feature (Gafni (¶0107 line 1-6), “electronic device 110-1 may ask the user for feedback on the determined one or more inspection criteria (which are, in essence, recommended inspection criteria). For example, the user may be asked to confirm or accept a recommended inspection criterion, or to modify or change it”)
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify image of Zueck in view of Wright instead be an image analysis module taught by Gafni, with a reasonable expectation of success. The motivation would be to “allow a user to interactively specify the information needed to simply and efficiently generate and distribute an application that leverage image acquisition capabilities of the second electronic devices, image analysis and logical rules (the one or more inspection criteria)” (Gafni (¶0049 line 1-6)).
As Claim 11, besides Claim 9, Zueck in view of Wright in further view of Gafni teaches wherein the operations further comprise:
training the machine learning model to identify the particular feature in images (Gafni (¶0114 line 5-10), “multiple images may be used in one or more ways. For example, multiple images of a scene may be acquired from different points of view in order to train object detection and classification”) and to activate the audio recording function in response to (Zueck (col. 6 line 36-40, fig. 6), “At block 206, the switch22 is activated that triggers a voice generated picture file mode after the user activates the switch. In one illustrative embodiment the switch automatically activates various steps associated with the voice generated picture file mode”) identifying the particular feature (Gafni (¶0107 line 1-6), “electronic device 110-1 may ask the user for feedback on the determined one or more inspection criteria (which are, in essence, recommended inspection criteria). For example, the user may be asked to confirm or accept a recommended inspection criterion, or to modify or change it”).
As Claim 13, besides Claim 1, Zueck in view of Wright may not explicitly disclose:
wherein detecting one or more features of the image comprises:
identifying, via an image analysis operation, an image feature in the image, wherein the image analysis operation comprises generating a feature vector based on image data of the image, and identifying the image feature based on the feature vector.
Gafni teaches:
wherein detecting one or more features of the image comprises:
identifying, via an image analysis operation, an image feature in the image (Gafni (¶0107 line 1-6), “electronic device 110-1 may ask the user for feedback on the determined one or more inspection criteria (which are, in essence, recommended inspection criteria). For example, the user may be asked to confirm or accept a recommended inspection criterion, or to modify or change it”), wherein the image analysis operation comprises generating a feature vector based on image data of the image, and identifying the image feature based on the feature vector (Gafni (¶0114 last 4 lines), “evidence may be integrated across multiple frames, which may allow a state vector (feature vector) that localizes objects and the camera for each frame to be probabilistically updated.”)
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify image of Zueck in view of Wright instead be an image analysis module taught by Gafni, with a reasonable expectation of success. The motivation would be to “allow a user to interactively specify the information needed to simply and efficiently generate and distribute an application that leverage image acquisition capabilities of the second electronic devices, image analysis and logical rules (the one or more inspection criteria)” (Gafni (¶0049 line 1-6)).
Claim(s) 7 and 10 is/are rejected under 35 U.S.C. 103 as being unpatentable over Zueck and Wright in view of Gafni in further view of Sikorski (U.S. 2004/0245334 hereinafter Sikorski).
As Claim 7, besides Claim 4, Zueck in view of Wright in further view of Gafni in further view of Arnold may not explicitly disclose:
wherein executing the machine learning model further comprises:
responsive to determining that the image meets the similarity threshold, selecting a return process application for execution by the mobile computing device to initiate a return of a damaged product, wherein the image meeting the similarity threshold indicates that the image depicts the damaged product
Sikorski teaches:
wherein executing the machine learning model further comprises:
responsive to determining that the image meets the similarity threshold, selecting a return process application for execution by the mobile computing device to initiate a return of a damaged product (Sikorski (¶0010 line 10-12), “customer decides to return one of the items previously added to the purchased item list, he scans the item bar code using the minus trigger key”), wherein the image meeting the similarity threshold indicates that the image depicts the damaged product (Sikorski (¶0048 line 8-12), “Once the damaged good is determined, the scanning mobile device can act ( e.g., order new good, request guidance, inform manufacturer, …) upon user and/or appointed authority ( e.g., artificial intelligence technique)”).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify image analysis of Zueck in view of Wright in further view of Gafni in further view of Arnold instead be an image analysis module taught by Sikorski, with a reasonable expectation of success. The motivation would be so that “a user can capture, analyze, and/or determine product identity based at least in part upon an image of a damaged good” (Sikorski (¶0048 line 6-8)).
As Claim 10, besides Claim 9, Zueck in view of Wright in further view of Gafni in teaches:
wherein the audio recording comprises a description of the damaged product (Zueck (col. 6 line 36-40, fig. 6), “At block 206, the switch22 is activated that triggers a voice generated picture file mode after the user activates the switch. In one illustrative embodiment the switch automatically activates various steps associated with the voice generated picture file mode”).
Zueck in view of Wright in further view of Gafni may not explicitly disclose:
wherein the particular feature comprises a damaged product
Sikorski teaches:
wherein the particular feature comprises a damaged product (Sikorski (¶0048 line 8-12), “Once the damaged good is determined, the scanning mobile device can act ( e.g., order new good, request guidance, inform manufacturer, …) upon user and/or appointed authority ( e.g., artificial intelligence technique)”).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify image analysis of Zueck in view of Wright in further view of Gafni in further view of Arnold instead be an image analysis module taught by Sikorski, with a reasonable expectation of success. The motivation would be so that “a user can capture, analyze, and/or determine product identity based at least in part upon an image of a damaged good” (Sikorski (¶0048 line 6-8)).
Response to Arguments
Rejections under 35 U.S.C. §101:
As Claim 1, Applicant argues that human mind cannot detect the capture of an image by an image capture function (second paragraph of page 8 in the remarks).
Applicant’s arguments are fully considered but are not persuasive. Human mind can detect a trigger/change from one state to another. Therefore, human mind can detect “the capture of an image”. For example, when a new image is displayed, human can understand that is the capture of an image.
As Claim 1, Applicant argues that human mind cannot features of an image (third paragraph of page 8 in the remarks).
Applicant’s arguments are fully considered but are not persuasive. Human mind can detect size shape and recognize images.
As Claim 1, Applicant argues that human mind cannot cause activation function to capture an audio recording (fourth paragraph of page 8 in the remarks).
Applicant’s arguments are fully considered but are not persuasive. Human mind can make a decision (such as “based on … “). The limitation is merely based on human mind to activate a function. The “activation …” is merely instruction to apply an exception (see MPEP §2106.05(f)). The limitation is analyzed together as one in order to treat the whole limitation instead of breaking the limitation into pieces.
As Claim 1, Applicant argues that the Claim requires a processor to execute the operations (last paragraph of page 8 in the remarks).
Applicant’s arguments are fully considered but are not persuasive. The processor execution steps are construed as mere instructions to apply an exception (see MPEP §2106.05(f)).
As Claim 1, Applicant argues that the claimed invention would “improving data storage operations of the mobile computing device” (last paragraph of page 10 in the remarks).
Applicant’s arguments are fully considered but are not persuasive. The “improving data storage operations” is not discussed or suggest by the Applicant’s specification.
As Claim 1, Applicant argues that the claimed invention would “avoiding ‘the inconvenience of switching between different computing applications on the mobile device’ ¶[0128]” (last paragraph of page 10 in the remarks).
Applicant’s arguments are persuasive; therefore, 35 U.S.C. §101 rejections are respectfully withdrawn. And further arguments regarding 35 U.S.C. §101 rejections are moot.
Rejections under 35 U.S.C. §102 and §103:
Applicant argues that Zueck fails to disclose “storing audio recording as an audio file” and “activating an audio recording function to capture an audio recording responsive to detecting a storage of an image” (last 2 paragraph of page 15 in the remarks).
Regarding “storing audio recording as an audio file”, applicant’s arguments are moot because new reference Wright teaches the limitation.
Regarding “activating an audio recording function …”, Zueck (col. 6 line 36-47 and 54-62, fig. 6 item 208), “At block 206, the switch 22 is activated that triggers a voice generated picture file mode after the user activates the switch. In one illustrative embodiment the switch automatically activates various steps associated with the voice generated picture file mode”, “In one embodiment, the switch 22 operations at block 206 are bypassed and the method proceeds to block 208” and “Typically the processing of the voice message occurs when an image or video is captured by the camera 26” teaches that the switch is bypassed. However, audio is automatically activated when the image was captured by the camera.
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure. Slaverry et al. (U.S. 2014/0164927) discloses a system/method to associate talk tags with captured image.
Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a).
A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to NHAT HUY T NGUYEN whose telephone number is (571)270-7333. The examiner can normally be reached M-F: 12:00-8:00 EST.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Viker Lamardo can be reached at 571-270-5871. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/NHAT HUY T NGUYEN/ Primary Examiner, Art Unit 2147