DETAILED ACTION
This non-final rejection is responsive to the claims filed 20 November 2024. Claims 1-20 are pending. Claims 1, 12, and 20 are independent claims.
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Claim Rejections - 35 USC § 102
The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action:
A person shall be entitled to a patent unless –
(a)(1) the claimed invention was patented, described in a printed publication, or in public use, on sale, or otherwise available to the public before the effective filing date of the claimed invention.
(a)(2) the claimed invention was described in a patent issued under section 151, or in an application for patent published or deemed published under section 122(b), in which the patent or application, as the case may be, names another inventor and was effectively filed before the effective filing date of the claimed invention.
Claims 1, 12, and 20 are rejected under 35 U.S.C. 102(a)(1) and (a)(2) as being anticipated by Fleizach (US 2010/0313125 A1) hereinafter known as Fleizach.
Regarding independent claim 1, Fleizach teaches:
displaying a target interface, wherein at least one interface object is displayed on the target interface, and the interface object comprises an interface display resource and/or an interface display control; (Fleizach: Fig. 5A and ¶[0346]; Fleizach teaches a first screen with interface elements consisting off menu items, dialog boxes, buttons, etc...)
acquiring, in response to an object trigger operation being input on an interface object, the triggered interface object as a target object; and (Fleizach: ¶[0347]-¶[0349]; Fleizach teaches a first finger gesture that corresponds to any of the plurality of user interface elements.)
determining voice introduction information corresponding to the target object; and (Fleizach: ¶[0349]-¶[0350]; Fleizach teaches that in response to detecting the finger gesture, outputting accessibility information, via audible information, associated with the second user interface element.)
playing the voice introduction information. (Fleizach: ¶[0349]-¶[0351]; Fleizach teaches emitting accessibility information as spoken text that corresponds to the second user interface element.)
Regarding independent claims 12 and 20, these claims recite an electronic device and a non-transitory storage medium that performs the method of claim 1; therefore, the same rationale for rejection applies. Fleizach further teaches processors and storage. (Fleizach: ¶[0068])
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claims 2, 3, 13, and 14 are rejected under 35 U.S.C. 103 as being unpatentable over Fleizach in view of Wu (US 2018/0181832 A1) hereinafter known as Wu.
Regarding claim 2, Fleizach further teaches the interface interaction method according to claim 1.
Fleizach does not explicitly teach but Wu teaches:
wherein determining the voice introduction information corresponding to the target object comprises: generating, in response to the target object being the interface display resource, the voice introduction information corresponding to the target object based on resource content of the interface display resource. (Wu: ¶[0026]-¶[0032]; Wu teaches generating descriptions of images using concepts detect by machine learning.
Fleizach and Wu are in the same field of endeavor as the present invention, as the references are directed to generating accessibility information. It would have been obvious, before the effective filing date of the claimed invention, to a person of ordinary skill in the art, to combine a system for generating accessibility spoken prompt for user selected interface elements as taught in Fleizach with generating the content based on resource content as taught in Wu. Fleizach already teaches speaking resource definitions. (Fleizach: ¶[0028]) Wu provides the additional functionality of generating the content based on the resource content As such, it would have been obvious to one of ordinary skill in the art to modify the teachings of Fleizach to include teachings of Wu, because the combination would allow reading the image descriptions to help the visually impaired, as suggested by Wu: ¶[0046].
Regarding claim 3, Fleizach in view of Wu further teaches the interface interaction method according to claim 2.
Wu further teaches:
wherein generating the voice introduction information corresponding to the target object based on the resource content of the target object comprises: generating an object keyword corresponding to the target object based on the resource content of the target object; and (Wu: ¶[0036]; Wu teaches an object recognition module to identify one or more object concepts depicted in an image.)
generating, based on the object keyword and preset description prompt information, a content introduction text corresponding to the target object, and converting the content introduction text into the voice introduction information. (Wu: ¶[0045]-¶[0046]; Wu teaches generating an image description based on concepts identified in the image (object keyword) and based on predetermined order based on concept type (preset description prompt information). Further, ¶[0046] teaches a screen reader reading the image description.)
Regarding claims 13 and 14, these claims recite an electronic device that performs the method of claims 2 and 3; therefore, the same rationale for rejection applies.
Claims 4 and 15 are rejected under 35 U.S.C. 103 as being unpatentable over Fleizach in view of Wu in view of Beshara (US 2023/0237280 A1) hereinafter known as Beshara.
Regarding claim 4, Fleizach in view of Wu further teaches the interface interaction method according to claim 3.
Wu further teaches:
wherein the interface display resource comprises an image resource, and generating the object keyword corresponding to the target object based on the resource content of the target object comprises: inputting the image resource to a content recognition model for content recognition to obtain the object keyword corresponding to the image resource, (Wu: ¶[0036]; Wu teaches an object recognition module to identify one or more object concepts depicted in an image.)
…
Fleizach in view of Wu does not explicitly teach but Beshara teaches:
wherein the content recognition model is obtained by training a neural network model based on a sample image and an expected keyword corresponding to the sample image, and the expected keyword is a keyword associated with image content of the sample image. (Beshara: ¶[0034] and ¶[0055]; Beshara teaches training AI models using annotated images.)
Beshara is in the same field of endeavor as the present invention, since it is directed to generating accessibility information. It would have been obvious, before the effective filing date of the claimed invention, to a person of ordinary skill in the art, to combine a system for generating accessibility spoken prompt for user selected interface elements while submitting an image to a machine learning model to generate a keyword as taught in Fleizach in view of Wu with further training an AI model using annotated images as taught in Beshara. As such, it would have been obvious to one of ordinary skill in the art to modify the teachings of Fleizach to include teachings of Beshara, because the combination would allow improved accessibility, as suggested by Beshara: ¶[0030].
Regarding claim 15, this claim recites an electronic device that performs the method of claim 4; therefore, the same rationale for rejection applies.
Claims 5 and 16 are rejected under 35 U.S.C. 103 as being unpatentable over Fleizach in view of Wu in view of Mahyar (US 10,999,566 B1) hereinafter known as Mahyar.
Regarding claim 5, Fleizach in view of Wu further teaches the interface interaction method according to claim 3.
Fleizach in view of Wu does not explicitly teach but Mahyar further teaches:
wherein the interface display resource comprises a video resource, and generating the object keyword corresponding to the target object based on the resource content of the target object comprises: acquiring a plurality of key frames in the video resource, and respectively performing content recognition on each key frame to obtain frame content keywords; and determining the object keyword corresponding to the video resource based on association relationships among the plurality of key frames and the frame content keywords corresponding to the key frames. (Mahyar: col. 3, lines 65-67 to col. 4, lines 1-21; Mahyar teaches a textual description generation engine that extracts the frames from the video content. Col. 15, lines 28-57 further teaches extracting frames at every 5 seconds of content. Col. 6, lines 12-25 and col. 4, lines 22-35 further teach detecting objects in the frames along with actions, i.e. standing/driving.)
Mahyar is in the same field of endeavor as the present invention, since it is directed to generating accessibility information. It would have been obvious, before the effective filing date of the claimed invention, to a person of ordinary skill in the art, to combine a system for generating accessibility spoken prompt for user selected interface elements while submitting an image to a machine learning model to generate a keyword as taught in Fleizach in view of Wu with further generating the object keyword by analyzing frames of a video as taught in Mahyar. As such, it would have been obvious to one of ordinary skill in the art to modify the teachings of Fleizach to include teachings of Mahyar, because the combination would allow improved accessibility, as suggested by Mahyar: col. 3, lines 8-14.
Regarding claim 16, this claim recites an electronic device that performs the method of claim 5; therefore, the same rationale for rejection applies.
Claims 6 and 17 are rejected under 35 U.S.C. 103 as being unpatentable over Fleizach in view of Wu in view of Dai (US 2024/0160858 A1) hereinafter known as Dai
Regarding claim 6, Fleizach in view of Wu further teaches the interface interaction method according to claim 3.
Wu further teaches:
wherein generating, based on the object keyword and the preset description prompt information, the content introduction text corresponding to the target object comprises: inputting the object keyword and the preset description prompt information to a text generation model to generate the content introduction text corresponding to the target object, … . (Wu: ¶[0045]-¶[0046]; Wu teaches generating an image description based on concepts identified in the image (object keyword) and based on predetermined order based on concept type (preset description prompt information). Further, ¶[0046] teaches a screen reader reading the image description.)
Fleizach in view of Wu does not explicitly teach but Beshara teaches:
… wherein the text generation model is obtained by training a deep learning model based on sample keywords, sample prompt information, and an expected introduction text. (Beshara: ¶[0045]; Beshara teaches fine tuning with additional data with expected output and text.)
Beshara is in the same field of endeavor as the present invention, since it is directed to generating accessibility information. It would have been obvious, before the effective filing date of the claimed invention, to a person of ordinary skill in the art, to combine a system for generating accessibility spoken prompt for user selected interface elements using a model as taught in Fleizach in view of Wu with further training an AI model using text and expected output as taught in Beshara. As such, it would have been obvious to one of ordinary skill in the art to modify the teachings of Fleizach to include teachings of Beshara, because the combination would allow improved accessibility, as suggested by Beshara: ¶[0030].
Fleizach in view of Wu in view of Beshara does not explicitly teach but Dai teaches:
… sample prompt information ... (Dai: ¶[0038]-¶[0040]; Dai teaches training ground truth input/output pairs, including an input image, an instruction, and a known good output text.)
Dai is analogous to the present invention, since it is reasonably pertinent to the problem faced by the inventor, i.e. generating output text. It would have been obvious, before the effective filing date of the claimed invention, to a person of ordinary skill in the art, to combine a system for generating accessibility spoken prompt for user selected interface elements using a model with training an AI model using text and expected output as taught in Fleizach in view of Wu in view of Beshara with further training based on sample prompt information as taught in Dai. As such, it would have been obvious to one of ordinary skill in the art to modify the teachings of Fleizach to include teachings of Dai, because the combination would allow efficiently representing aspects of the image, as suggested by Dai: ¶[0038].
Regarding claim 17, this claim recites an electronic device that performs the method of claim 6; therefore, the same rationale for rejection applies.
Claims 7, 8, 10, 18, and 19 are rejected under 35 U.S.C. 103 as being unpatentable over Fleizach.
Regarding claim 7, Fleizach further teaches the interface interaction method according to claim 1.
Another embodiment of Fleizach teaches:
wherein determining the voice introduction information corresponding to the target object comprises: generating, in response to the target object being the interface display control, the voice introduction information corresponding to the interface display control based on function associated information corresponding to the interface display control. (Fleizach: ¶[0352]; Fleizach teaches generating spoken text, such as volume control – swipe up to increase.)
Fleizach is in the same field of endeavor as the present invention, as it is directed to generating accessibility information. It would have been obvious, before the effective filing date of the claimed invention, to a person of ordinary skill in the art, to combine a system for generating accessibility spoken prompt for user selected interface elements as taught in Fleizach with further generating spoken text for display controls. As such, it would have been obvious to one of ordinary skill in the art to combine these teachings because the combination would allow notifying the user of their functions, as suggested by Fleizach: ¶[0352].
Regarding claim 8, Fleizach further teaches the interface interaction method according to claim 7.
Fleizach further teaches:
wherein the function associated information comprises an acting object corresponding to the interface display control and an acting result generated after the interface display control acts on the acting object; and generating the voice introduction information corresponding to the interface display control based on the function associated information corresponding to the interface display control comprises: determining the acting object corresponding to the interface display control and the acting result generated after the interface display control acts on the acting object; generating a function description text corresponding to the interface display control based on the acting object and the acting result, and converting the function description text into the voice introduction information. (Fleizach: ¶[0352]; Fleizach teaches generating spoken text, such as volume control – swipe up to increase. ¶[0249]-¶[0251] and ¶[0350]-¶[0351] further teaches announcing the changed value.)
Regarding claim 10, Fleizach further teaches the interface interaction method according to claim 1.
Fleizach further teaches:
wherein acquiring, in response to the object trigger operation being input on the interface object, the triggered interface object as the target object comprises: determining, in response to a touch selection operation being input on the interface object, the interface object selected based on the touch selection operation as the target object; and/or, determining, in response to an object gaze operation being input on the interface object, a gaze fixation area, and determining the interface object corresponding to the gaze fixation area as the target object. (Fleizach: ¶[0201]-¶[0202]; Fleizach teaches a single finger tap on a music application icon, which produces audible information.)
Fleizach is in the same field of endeavor as the present invention, as it is directed to generating accessibility information. It would have been obvious, before the effective filing date of the claimed invention, to a person of ordinary skill in the art, to combine a system for generating accessibility spoken prompt for user selected interface elements as taught in Fleizach with further generating spoken text in response to a touch selection. As such, it would have been obvious to one of ordinary skill in the art to combine these teachings because the combination would allow notifying the user of their functions, as suggested by Fleizach: ¶[0352].
Regarding claims 18 and 19, these claims recite an electronic device that performs the method of claims 7 and 8; therefore, the same rationale for rejection applies.
Claim 9 is rejected under 35 U.S.C. 103 as being unpatentable over Fleizach in view of Reiter (US 2014/0067377 A1) hereinafter known as Reiter.
Regarding claim 9, Fleizach further teaches the interface interaction method according to claim 8.
Fleizach does not explicitly teach but Reiter teaches:
wherein generating the function description text corresponding to the interface display control based on the acting object and the acting result comprises: generating a control keyword corresponding to the interface display control based on the acting object and the acting result; and generating the function description text corresponding to the interface display control based on the control keyword and the preset description prompt information. (Reiter: ¶[0039]-¶[0040] and ¶[0068]; Reiter teaches generating one or more messages that are populated or otherwise instantiated based on data or information in the primary data channel – corresponding to a fact about the underlying data. Further, ¶[0059] and ¶[0070]-¶[0071] teach a rule set that defines the order in which a number of message are to be presented in a document and syntactic constituents and features and outputting a well-formed natural language text.)
Reiter is analogous to the present invention, since it is reasonably pertinent to the problem faced by the inventor, i.e. generating output text. It would have been obvious, before the effective filing date of the claimed invention, to a person of ordinary skill in the art, to combine a system for generating accessibility spoken prompt for user selected interface elements using a model with training an AI model using text and expected output as taught in Fleizach with generating the content based a keyword and preset description prompt information as taught in Reiter. As such, it would have been obvious to one of ordinary skill in the art to modify the teachings of Fleizach to include teachings of Reiter, because the combination would allow producing situational analysis text for alert conditions, as suggested by Reiter: ¶[0016].
Claim 11 is rejected under 35 U.S.C. 103 as being unpatentable over Fleizach in view of Beshara.
Regarding claim 11, Fleizach further teaches the interface interaction method according to claim 1.
Fleizach further teaches:
wherein determining the voice introduction information corresponding to the target object comprises: acquiring, in response to detecting text introduction information corresponding to the target object, the text introduction information, and converting the text introduction information into the voice introduction information; and (Fleizach: ¶[0228] and ¶[0368]; Fleizach teaches detecting and speaking descriptions of images.)
…
Fleizach does not explicitly teach but Beshara teaches:
generating the voice introduction information based on the target object in response to not detecting the text introduction information corresponding to the target object. (Beshara: ¶[0037]; Beshara teaches making a determination that an image is missing ALT text and processing the image by an image captioning model.)
Beshara is in the same field of endeavor as the present invention, since it is directed to generating accessibility information. It would have been obvious, before the effective filing date of the claimed invention, to a person of ordinary skill in the art, to combine a system for generating accessibility spoken prompt for user selected interface elements as taught in Fleizach with generating missing text using an image captioning model as taught in Beshara. As such, it would have been obvious to one of ordinary skill in the art to modify the teachings of Fleizach to include teachings of Beshara, because the combination would allow improved accessibility, as suggested by Beshara: ¶[0030].
Conclusion
Any inquiry concerning this communication or earlier communications from the examiner should be directed to ALEX OLSHANNIKOV whose telephone number is (571)270-0667. The examiner can normally be reached M-F 9:30-6.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Scott Baderman can be reached at 571-272-3644. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/ALEKSEY OLSHANNIKOV/Primary Examiner, Art Unit 2118