Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
DETAILED ACTION
This office action is in response to correspondence filed 06/29/26 regarding application 18/886,702, in which claims 1, 4, 5, 7, 9, 11, 13, 16, 17, 19, and 20 were amended and claims 3 and 15 were cancelled. Claims 1, 2, 4-14, and 16-20 are pending and have been considered.
Response to Arguments
The examiner agrees with Applicant on page 15 that the amendments to claims 1, 4, 5, 7, 9, 11, 13, 16, 17, 19, and 20 are fully supported by the original disclosure. The examiner finds that the amendments contain no new matter.
The examiner agrees with Applicant on pages 15-16 that as amended, the claims are not directed to an abstract idea in the form of a mental process but rather performing speech recognition using a language model parameter adjusted based on screen text information. The 35 U.S.C. 101 rejections of claims 1, 2, 4-14, and 16-20 are withdrawn.
Applicant’s arguments on pages 19-22 regarding the 35 U.S.C. 35 U.S.C. 103 rejections of claims 1, 2, 4-14, and 16-20 based on Corcoran, Fu, and Manjunath have been considered but are moot in view of the new grounds for rejection based in part on the newly discovered reference to Sarukkai et al. (US 5819220), which is considered very relevant to the “… the at least one piece of second text information comprises at least one word not included in the first text information…” language of amended claims 1, 13, and 20 because Sarukkai describes a set of words is extracted from the web page source that is currently being displayed by a browser, and then a different set of words “related” to the source word set is determined for word set probability boosting in the language model for speech recognition, see Col 7-8 lines 63-11. The new grounds for rejection based in part on Sarukkai are necessitated by the claim amendments.
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102 of this title, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claims 1, 2, 4, 5, 7-14, 16, 17, 19, and 20 are rejected under 35 U.S.C. 103 as being unpatentable over Corcoran et al. (US 20180211669 A1) in view of Fu (CN 113838460 A), in further view of Sarukkai et al. (US 5819220).
Consider claim 1, Corcoran discloses a discloses a speech-to-text conversion method, performed by an electronic device (converting speech to text on an electronic device, [0027]-[0028]), the method comprising:
obtaining speech information to be converted and a screen (voice commands received as audio input via microphone, [0018], to be converted into text, [0022], and a foreground window representing a piece of software displayed, i.e. a screen, [0025]), the screen being related to the speech information and displayed on a screen during generation of the speech information (a window currently displayed in the foreground when the audio is received, [0025]);
performing speech-to-text conversion on the speech information to obtain a plurality of pieces of candidate text information (a version of the audio is sent to a plurality of speech recognition engines, which each return a text version of the audio, i.e. candidate text information, [0025]);
determining screen text information based on the screen, wherein the screen text information comprises first text information obtained from the screen image (choices available on the screen in the present foreground window, [0026]);
determining target appearance indicators corresponding to the plurality of pieces of candidate text information based on the screen text information, a target appearance indicator of one piece of the candidate text information representing a probability that the one piece of candidate text information corresponds to the speech information (choices available on the screen in the present foreground window, i.e. indications that the words are more likely to be what the user said, [0026], [0009]); and
selecting the candidate text information whose target appearance indicator meets a requirement from the plurality of pieces of candidate text information as converted text information of the speech information (choosing a speech recognition to use by seeing if the text returned by the engines matches the choices available on the screen in the present foreground window, [0026]);
transmitting the converted text information to a terminal device for display on a screen (speech engines return a version of audio in text format, which is defined as a string of letters representative of the audio spoken which can be shown, i.e. displayed, in human readable format on window displayed on screen of computing device, [0022], [0023], the examiner noting that the claim does not actually require displaying text on the screen of the terminal device, but merely that the text has an intended use of “for display on a screen).
Corcoran does not specifically mention a screen image;
determining at least one piece of second text information generated based on the first text information, and semantics of the second text information is the same as, opposite to, or similar to semantics of the first text information;
adjusting a parameter of a language model based on the screen text information to obtain an adjusted language model, the language model having been trained to learn a probability that a word appears in a general scenario, and the adjusted language model learning a probability that the word appears in an information sharing scenario reflected by the screen text information;
determining, by invoking the adjusted language model, target appearance indicators.
Fu discloses a screen image (image frame sequence, page 4);
determining at least one piece of second text information generated based on the first text information, and semantics of the second text information is the same as, opposite to, or similar to semantics of the first text information (the words are filtered to produce key words, page 10; since these are filtered words, they have the same semantics as the corresponding words recognized from the text);
adjusting a parameter of a language model based on the screen text information to obtain an adjusted language model (statistical probability of text, i.e. base probability in the language model, is augmented with contribution probability of keyword in text b, page 8), the language model having been trained to learn a probability that a word appears in a general scenario, and the adjusted language model learning a probability that the word appears in an information sharing scenario reflected by the screen text information (statistical probability of text is the base language model probability for a general scenario, for which training is inherent to determine the base probabilities, and the adjusted probability augmented by contribution probability of keyword in text b is a probability that the word appears, considered a learned probability, page 8).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the invention of Corcoran by including a screen image, determining at least one piece of second text information generated based on the first text information, and semantics of the second text information is the same as, opposite to, or similar to semantics of the first text information, adjusting a parameter of a language model based on the screen text information to obtain an adjusted language model, the language model having been trained to learn a probability that a word appears in a general scenario, and the adjusted language model learning a probability that the word appears in an information sharing scenario reflected by the screen text information, and determining as in Corcoran, by invoking the adjusted language model of Fu, target appearance indicators as in Corcoran in order to improve voice recognition accuracy, as suggested by Fu (page 4). Doing so would have led to predictable results of mitigating the adverse effects of background noise and other interference, as suggested by Fu (page 2). The references cited are analogous art in the same field of speech recognition.
Corcoran and Fu do not specifically mention the at least one piece of second text information comprises at least one word not included in the first text information.
Sarukkai discloses the at least one piece of second text information comprises at least one word not included in the first text information (a set of words is extracted from the web page source that is currently being displayed by the browser, and then set of words “related” to the source word set is determined for word set boosting in the language model, Col 7-8 lines 63-11).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the invention of Corcoran and Fu such that the at least one piece of second text information comprises at least one word not included in the first text information in order to make the language model more dynamic, as suggested by Sarukkai (Col 2 lines 35-50), predictably overcoming the limitations of static language modeling described by Sarukkai (Col 2 lines 35-50), while also helping alleviate the well known out-of-vocabulary problem described by Sarukkai (Col 2 lines 1-16).
The references cited are analogous art in the same field of speech recognition.
Consider claim 13, Corcoran discloses a speech-to-text conversion apparatus (device used to carry out the disclosed technology by converting speech to text, [0027]-[0028]), comprising:
a processor and a memory, the memory having at least one computer program stored therein, and the at least one computer program being loaded and executed by the processor (program instructions loaded into memory and executed by processor of device, [0028]) to cause the electronic device to implement:
obtaining speech information to be converted and a screen (voice commands received as audio input via microphone, [0018], to be converted into text, [0022], and a foreground window representing a piece of software displayed, i.e. a screen, [0025]), the screen being related to the speech information and displayed on a screen during generation of the speech information (a window currently displayed in the foreground when the audio is received, [0025]);
performing speech-to-text conversion on the speech information to obtain a plurality of pieces of candidate text information (a version of the audio is sent to a plurality of speech recognition engines, which each return a text version of the audio, i.e. candidate text information, [0025]);
determining screen text information based on the screen, wherein the screen text information comprises first text information obtained from the screen image (choices available on the screen in the present foreground window, [0026]);
determining target appearance indicators corresponding to the plurality of pieces of candidate text information based on the screen text information, a target appearance indicator of one piece of the candidate text information representing a probability that the one piece of candidate text information corresponds to the speech information (choices available on the screen in the present foreground window, i.e. indications that the words are more likely to be what the user said, [0026], [0009]); and
selecting the candidate text information whose target appearance indicator meets a requirement from the plurality of pieces of candidate text information as converted text information of the speech information (choosing a speech recognition to use by seeing if the text returned by the engines matches the choices available on the screen in the present foreground window, [0026]);
transmitting the converted text information to a terminal device for display on a screen (speech engines return a version of audio in text format, which is defined as a string of letters representative of the audio spoken which can be shown, i.e. displayed, in human readable format on window displayed on screen of computing device, [0022], [0023], the examiner noting that the claim does not actually require displaying text on the screen of the terminal device, but merely that the text has an intended use of “for display on a screen).
Corcoran does not specifically mention a screen image;
determining at least one piece of second text information generated based on the first text information, and semantics of the second text information is the same as, opposite to, or similar to semantics of the first text information;
adjusting a parameter of a language model based on the screen text information to obtain an adjusted language model, the language model having been trained to learn a probability that a word appears in a general scenario, and the adjusted language model learning a probability that the word appears in an information sharing scenario reflected by the screen text information;
determining, by invoking the adjusted language model, target appearance indicators.
Fu discloses a screen image (image frame sequence, page 4);
determining at least one piece of second text information generated based on the first text information, and semantics of the second text information is the same as, opposite to, or similar to semantics of the first text information (the words are filtered to produce key words, page 10; since these are filtered words, they have the same semantics as the corresponding words recognized from the text);
adjusting a parameter of a language model based on the screen text information to obtain an adjusted language model (statistical probability of text, i.e. base probability in the language model, is augmented with contribution probability of keyword in text b, page 8), the language model having been trained to learn a probability that a word appears in a general scenario, and the adjusted language model learning a probability that the word appears in an information sharing scenario reflected by the screen text information (statistical probability of text is the base language model probability for a general scenario, for which training is inherent to determine the base probabilities, and the adjusted probability augmented by contribution probability of keyword in text b is a probability that the word appears, considered a learned probability, page 8).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the invention of Corcoran by including a screen image, determining at least one piece of second text information generated based on the first text information, and semantics of the second text information is the same as, opposite to, or similar to semantics of the first text information, adjusting a parameter of a language model based on the screen text information to obtain an adjusted language model, the language model having been trained to learn a probability that a word appears in a general scenario, and the adjusted language model learning a probability that the word appears in an information sharing scenario reflected by the screen text information, and determining as in Corcoran, by invoking the adjusted language model of Fu, target appearance indicators as in Corcoran for reasons similar to those for claim 1.
Corcoran and Fu do not specifically mention the at least one piece of second text information comprises at least one word not included in the first text information.
Sarukkai discloses the at least one piece of second text information comprises at least one word not included in the first text information (a set of words is extracted from the web page source that is currently being displayed by the browser, and then set of words “related” to the source word set is determined for word set boosting in the language model, Col 7-8 lines 63-11).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the invention of Corcoran and Fu such that the at least one piece of second text information comprises at least one word not included in the first text information for reasons similar to those for claim 1.
Consider claim 20, Corcoran discloses a non-transitory computer-readable storage medium, having at least one computer program stored therein, the at least one computer program being loaded and executed by a processor of an electronic device (program instructions loaded into memory and executed by processor of device, [0028]), to cause the electronic device to implement:
obtaining speech information to be converted and a screen (voice commands received as audio input via microphone, [0018], to be converted into text, [0022], and a foreground window representing a piece of software displayed, i.e. a screen, [0025]), the screen being related to the speech information and displayed on a screen during generation of the speech information (a window currently displayed in the foreground when the audio is received, [0025]);
performing speech-to-text conversion on the speech information to obtain a plurality of pieces of candidate text information (a version of the audio is sent to a plurality of speech recognition engines, which each return a text version of the audio, i.e. candidate text information, [0025]);
determining screen text information based on the screen, wherein the screen text information comprises first text information obtained from the screen image (choices available on the screen in the present foreground window, [0026]);
determining target appearance indicators corresponding to the plurality of pieces of candidate text information based on the screen text information, a target appearance indicator of one piece of the candidate text information representing a probability that the one piece of candidate text information corresponds to the speech information (choices available on the screen in the present foreground window, i.e. indications that the words are more likely to be what the user said, [0026], [0009]); and
selecting the candidate text information whose target appearance indicator meets a requirement from the plurality of pieces of candidate text information as converted text information of the speech information (choosing a speech recognition to use by seeing if the text returned by the engines matches the choices available on the screen in the present foreground window, [0026]);
transmitting the converted text information to a terminal device for display on a screen (speech engines return a version of audio in text format, which is defined as a string of letters representative of the audio spoken which can be shown, i.e. displayed, in human readable format on window displayed on screen of computing device, [0022], [0023], the examiner noting that the claim does not actually require displaying text on the screen of the terminal device, but merely that the text has an intended use of “for display on a screen).
Corcoran does not specifically mention a screen image;
determining at least one piece of second text information generated based on the first text information, and semantics of the second text information is the same as, opposite to, or similar to semantics of the first text information;
adjusting a parameter of a language model based on the screen text information to obtain an adjusted language model, the language model having been trained to learn a probability that a word appears in a general scenario, and the adjusted language model learning a probability that the word appears in an information sharing scenario reflected by the screen text information;
determining, by invoking the adjusted language model, target appearance indicators.
Fu discloses a screen image (image frame sequence, page 4);
determining at least one piece of second text information generated based on the first text information, and semantics of the second text information is the same as, opposite to, or similar to semantics of the first text information (the words are filtered to produce key words, page 10; since these are filtered words, they have the same semantics as the corresponding words recognized from the text);
adjusting a parameter of a language model based on the screen text information to obtain an adjusted language model (statistical probability of text, i.e. base probability in the language model, is augmented with contribution probability of keyword in text b, page 8), the language model having been trained to learn a probability that a word appears in a general scenario, and the adjusted language model learning a probability that the word appears in an information sharing scenario reflected by the screen text information (statistical probability of text is the base language model probability for a general scenario, for which training is inherent to determine the base probabilities, and the adjusted probability augmented by contribution probability of keyword in text b is a probability that the word appears, considered a learned probability, page 8).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the invention of Corcoran by including a screen image, determining at least one piece of second text information generated based on the first text information, and semantics of the second text information is the same as, opposite to, or similar to semantics of the first text information, adjusting a parameter of a language model based on the screen text information to obtain an adjusted language model, the language model having been trained to learn a probability that a word appears in a general scenario, and the adjusted language model learning a probability that the word appears in an information sharing scenario reflected by the screen text information, and determining as in Corcoran, by invoking the adjusted language model of Fu, target appearance indicators as in Corcoran for reasons similar to those for claim 1.
Corcoran and Fu do not specifically mention the at least one piece of second text information comprises at least one word not included in the first text information.
Sarukkai discloses the at least one piece of second text information comprises at least one word not included in the first text information (a set of words is extracted from the web page source that is currently being displayed by the browser, and then set of words “related” to the source word set is determined for word set boosting in the language model, Col 7-8 lines 63-11).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the invention of Corcoran and Fu such that the at least one piece of second text information comprises at least one word not included in the first text information for reasons similar to those for claim 1.
Consider claim 2, Corcoran discloses the performing speech-to-text conversion on the speech information to obtain a plurality of pieces of candidate text information comprises: performing speech-to-text conversion on speech to obtain a character corresponding to the speech (e.g. the characters “p”, “u”, “p” etc. corresponding to the “puppy” segment, [0024]); and determining the plurality of pieces of candidate text information based on the characters corresponding to the speech (interpreting the speech as “puppy formula”, [0024]).
Corcoran does not specifically mention performing segmentation on the speech information to obtain a plurality of speech segments.
Fu discloses performing segmentation on the speech information to obtain a plurality of speech segments (audio sub-segments containing speech, page 4).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the invention of Corcoran by performing segmentation on the speech information to obtain a plurality of speech segments for reasons similar to those for claim 1.
Consider claim 4, Corcoran does not, but Fu discloses determining screen text information based on the screen image comprises: performing image segmentation on the screen image to obtain at least one of a text image and an object image, wherein the text image reflects a text displayed on the screen, and the object image reflects an object displayed on the screen (sampling a picture every fP frame with image subtitles, pages 6-7; this is considered to form segments of screen images to obtain text images of the subtitles, see Fig 3 depicting text characters and objects); and determining the screen text information based on at least one of the text image and the object image (OCR generates the text results based on the frame images making up the segment, page 7).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the invention of Corcoran by determining screen text information based on the screen image comprises: performing image segmentation on the screen image to obtain at least one of a text image and an object image, wherein the text image reflects a text displayed on the screen, and the object image reflects an object displayed on the screen; and determining the screen text information based on at least one of the text image and the object image for reasons similar to those for claim 1.
Consider claim 5, Corcoran does not, but Fu discloses first text information is obtained by performing text recognition on the text image (OCR identifies text from the image, to produce text result L1, page 7).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the invention of Corcoran such that first text information is obtained by performing text recognition on the text image for reasons similar to those for claim 1.
Consider claim 7, Corcoran does not, but Fu discloses determining the target appearance indicators corresponding to the plurality of pieces of candidate text information based on the screen text information comprises: performing word segmentation on a piece of candidate text information to obtain words comprised in the piece of candidate text information (OCR results and segmentation into word list, page 7); determining a target appearance indicator of each word comprised in the piece of candidate text information based on the screen text information, wherein the target appearance indicator of the word represents a probability that the word is comprised in the speech information (identification probability of each candidate text, page 8); and determining the target appearance indicator of the piece of candidate text information based on the target appearance indicators of the words comprised in the piece of candidate text information (determining the target recognition results based on the probability, page 8).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the invention of Corcoran such that determining the target appearance indicators corresponding to the plurality of pieces of candidate text information based on the screen text information comprises: performing word segmentation on a piece of candidate text information to obtain words comprised in the piece of candidate text information; determining a target appearance indicator of each word comprised in the piece of candidate text information based on the screen text information, wherein the target appearance indicator of the word represents a probability that the word is comprised in the speech information; and determining the target appearance indicator of the piece of candidate text information based on the target appearance indicators of the words comprised in the piece of candidate text information for reasons similar to those for claim 1.
Consider claim 8, Corcoran does not, but Fu discloses the determining a target appearance indicator of each word comprised in the candidate text information based on the screen text information comprises: obtaining an initial appearance indicator of each word, wherein the initial appearance indicator of the word represents a probability that the word appears in a general scenario (statistical probability of the text appearing, page 8); and adjusting the initial appearance indicator of the word based on the screen text information to obtain the target appearance indicator of the word (multiplying by the contribution probability and coefficients, page 8).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the invention of Corcoran such that determining a target appearance indicator of each word comprised in the candidate text information based on the screen text information comprises: obtaining an initial appearance indicator of each word, wherein the initial appearance indicator of the word represents a probability that the word appears in a general scenario and adjusting the initial appearance indicator of the word based on the screen text information to obtain the target appearance indicator of the word for reasons similar to those for claim 1.
Consider claim 9, Corcoran does not, but Fu discloses adjusting the initial appearance indicator of the word based on the screen text information to obtain the target appearance indicator of the word comprises: determining a first appearance indicator of the word based on the first text information and the at least one piece of second text information, wherein the first appearance indicator of the word represents a probability that the word appears in the first text information and the at least one piece of second text information (statistical probability of the text, page 8); and performing weighted calculation on the first appearance indicator of the word and the initial appearance indicator of the word to obtain the target appearance indicator of the word (combining the acoustic model and language model probabilities with weighting coefficients, page 8).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the invention of Corcoran such that adjusting the initial appearance indicator of the word based on the screen text information to obtain the target appearance indicator of the word comprises: determining a first appearance indicator of the word based on the first text information and the at least one piece of second text information, wherein the first appearance indicator of the word represents a probability that the word appears in the first text information and the at least one piece of second text information; and performing weighted calculation on the first appearance indicator of the word and the initial appearance indicator of the word to obtain the target appearance indicator of the word for reasons similar to those for claim 1.
Consider claim 10, Corcoran does not, but Fu discloses the screen text information comprises third text information and at least one piece of fourth text information (e.g. text result L3 and corresponding word from word list W3, page 7); the third text information is text information corresponding to the object image in the screen image, and semantics of the fourth text information is the same as, opposite to, or similar to semantics of the third text information (W3 being filtered word L3 having the same semantics, page 7); and the adjusting the initial appearance indicator of the word based on the screen text information to obtain the target appearance indicator of the word comprises: determining a second appearance indicator of the word based on the third text information and the at least one piece of fourth text information, wherein the second appearance indicator of the word represents a probability that the word appears in the third text information and the at least one piece of fourth text information (statistical probability of the text, page 8); and performing weighted calculation on the second appearance indicator of the word and the initial appearance indicator of the word to obtain the target appearance indicator of the word (combining the acoustic model and language model probabilities with weighting coefficients, page 8).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the invention of Corcoran such that the screen text information comprises third text information and at least one piece of fourth text information; the third text information is text information corresponding to the object image in the screen image, and semantics of the fourth text information is the same as, opposite to, or similar to semantics of the third text information; and the adjusting the initial appearance indicator of the word based on the screen text information to obtain the target appearance indicator of the word comprises: determining a second appearance indicator of the word based on the third text information and the at least one piece of fourth text information, wherein the second appearance indicator of the word represents a probability that the word appears in the third text information and the at least one piece of fourth text information; and performing weighted calculation on the second appearance indicator of the word and the initial appearance indicator of the word to obtain the target appearance indicator of the word for reasons similar to those for claim 1.
Consider claim 11, Corcoran does not, but Fu discloses the third text information is text information corresponding to the object image in the screen image, and semantics of the fourth text information is the same as, opposite to, or similar to semantics of the third text information (W2 being filtered word L2 having the same semantics, page 7); and the adjusting the initial appearance indicator of the word based on the screen text information to obtain the target appearance indicator of the word comprises: determining a third appearance indicator of the word based on the first text information, the at least one piece of second text information, the third text information, and the at least one piece of fourth text information, wherein the third appearance indicator of the word represents a probability that the word appears in the first text information, the at least one piece of second text information, the third text information, and the at least one piece of fourth text information (statistical probability of the text, page 8); and performing weighted calculation on the third appearance indicator of the word and the initial appearance indicator of the word to obtain the target appearance indicator of the word (combining the acoustic model and language model probabilities with weighting coefficients, page 8).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the invention of Corcoran such that the third text information is text information corresponding to the object image in the screen image, and semantics of the fourth text information is the same as, opposite to, or similar to semantics of the third text information; and the adjusting the initial appearance indicator of the word based on the screen text information to obtain the target appearance indicator of the word comprises: determining a third appearance indicator of the word based on the first text information, the at least one piece of second text information, the third text information, and the at least one piece of fourth text information, wherein the third appearance indicator of the word represents a probability that the word appears in the first text information, the at least one piece of second text information, the third text information, and the at least one piece of fourth text information; and performing weighted calculation on the third appearance indicator of the word and the initial appearance indicator of the word to obtain the target appearance indicator of the word for reasons similar to those for claim 1.
Consider claim 12, Corcoran does not, but Fu discloses determining the target appearance indicator of the candidate text information based on the target appearance indicators of the words comprised in the candidate text information comprises: obtaining a first appearance indicator of the candidate text information, wherein the first appearance indicator of the candidate text information represents a probability that the speech information is converted into the candidate text information (recognition probability of the acoustic model, page 8); determining a second appearance indicator of the candidate text information based on the target appearance indicators of the words, wherein the second appearance indicator of the candidate text information represents a probability that the candidate text information is determined based on the words (statistical probability of the text, page 8); and performing weighted processing on the first appearance indicator of the candidate text information and the second appearance indicator of the candidate text information to obtain the target appearance indicator of the candidate text information (combining the acoustic model and language model probabilities with weighting coefficients, page 8).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the invention of Corcoran such that determining the target appearance indicator of the candidate text information based on the target appearance indicators of the words comprised in the candidate text information comprises: obtaining a first appearance indicator of the candidate text information, wherein the first appearance indicator of the candidate text information represents a probability that the speech information is converted into the candidate text information determining a second appearance indicator of the candidate text information based on the target appearance indicators of the words, wherein the second appearance indicator of the candidate text information represents a probability that the candidate text information is determined based on the words and performing weighted processing on the first appearance indicator of the candidate text information and the second appearance indicator of the candidate text information to obtain the target appearance indicator of the candidate text information for reasons similar to those for claim 1.
Consider claim 14, Corcoran discloses the performing speech-to-text conversion on the speech information to obtain a plurality of pieces of candidate text information comprises: performing speech-to-text conversion on speech to obtain a character corresponding to the speech (e.g. the characters “p”, “u”, “p” etc. corresponding to the “puppy” segment, [0024]); and determining the plurality of pieces of candidate text information based on the characters corresponding to the speech (interpreting the speech as “puppy formula”, [0024]).
Corcoran does not specifically mention performing segmentation on the speech information to obtain a plurality of speech segments.
Fu discloses performing segmentation on the speech information to obtain a plurality of speech segments (audio sub-segments containing speech, page 4).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the invention of Corcoran by performing segmentation on the speech information to obtain a plurality of speech segments for reasons similar to those for claim 1.
Consider claim 16, Corcoran does not, but Fu discloses determining screen text information based on the screen image comprises: performing image segmentation on the screen image to obtain at least one of a text image and an object image, wherein the text image reflects a text displayed on the screen, and the object image reflects an object displayed on the screen (sampling a picture every fP frame with image subtitles, pages 6-7; this is considered to form segments of screen images to obtain text images of the subtitles, see Fig 3 depicting text characters and objects); and determining the screen text information based on at least one of the text image and the object image (OCR generates the text results based on the frame images making up the segment, page 7).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the invention of Corcoran by determining screen text information based on the screen image comprises: performing image segmentation on the screen image to obtain at least one of a text image and an object image, wherein the text image reflects a text displayed on the screen, and the object image reflects an object displayed on the screen; and determining the screen text information based on at least one of the text image and the object image for reasons similar to those for claim 1.
Consider claim 17, Corcoran does not, but Fu discloses first text information is obtained by performing text recognition on the text image (OCR identifies text from the image, to produce text result L1, page 7).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the invention of Corcoran such that first text information is obtained by performing text recognition on the text image for reasons similar to those for claim 1.
Consider claim 19, Corcoran does not, but Fu discloses determining the target appearance indicators corresponding to the plurality of pieces of candidate text information based on the screen text information comprises: performing word segmentation on a piece of candidate text information to obtain words comprised in the piece of candidate text information (OCR results and segmentation into word list, page 7); determining a target appearance indicator of each word comprised in the piece of candidate text information based on the screen text information, wherein the target appearance indicator of the word represents a probability that the word is comprised in the speech information (identification probability of each candidate text, page 8); and determining the target appearance indicator of the piece of candidate text information based on the target appearance indicators of the words comprised in the piece of candidate text information (determining the target recognition results based on the probability, page 8).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the invention of Corcoran such that determining the target appearance indicators corresponding to the plurality of pieces of candidate text information based on the screen text information comprises: performing word segmentation on a piece of candidate text information to obtain words comprised in the piece of candidate text information; determining a target appearance indicator of each word comprised in the piece of candidate text information based on the screen text information, wherein the target appearance indicator of the word represents a probability that the word is comprised in the speech information; and determining the target appearance indicator of the piece of candidate text information based on the target appearance indicators of the words comprised in the piece of candidate text information for reasons similar to those for claim 1.
Claims 6 and 18 are rejected under 35 U.S.C. 103 as being unpatentable over Corcoran et al. (US 20180211669 A1) in view of Fu (CN 113838460 A), in further view of Sarukkai et al. (US 5819220), in further view of Manjunath et al. (US 20210151038 A1).
Consider claim 6, Corcoran does not, but Fu discloses determining the screen text information based on the object image comprises: performing image processing on the object image to obtain third text information, (using OCR to perform text recognition on the caption, page 7); determining at least one piece of fourth text information based on the third text information, wherein semantics of the fourth text information is the same as, opposite to, or similar to semantics of the third text information (the words from the caption are filtered to produce key words, pages 7, 10; since these are filtered words, they have the same semantics as the corresponding words recognized from the text); and determining the third text information and the at least one piece of fourth text information as the screen text information (the recognized words of the caption selected as keywords for filtering are used as screen text information for recognition probabilities, page 10).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the invention of Corcoran such that determining the screen text information based on the object image comprises: performing image processing on the object image to obtain third text information,; determining at least one piece of fourth text information based on the third text information, wherein semantics of the fourth text information is the same as, opposite to, or similar to semantics of the third text information; and determining the third text information and the at least one piece of fourth text information as the screen text information for reasons similar to those for claim 1.
Corcoran, Fu, and Sarukkai do not specifically mention performing image description processing on the object image to obtain third text information, wherein the third text information describes the object in the object image.
Manjunath discloses performing image description processing on an object image to obtain text information, wherein the text information describes the object in the object image (extracting objects from the image frames and analyzing them to generate keywords indicating textual descriptors, [0051]).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the invention of Corcoran, Fu, and Sarukkai by performing image description processing on the object image to obtain third text information, wherein the third text information describes the object in the object image in order to compensate for variations in audio quality adversely affecting recognition confidence, as suggested by Manjunath ([0032]), predictably improving accuracy, as suggested by Manjunath ([0058]). The references cited are analogous art in the same field of speech recognition.
Consider claim 18, Corcoran does not, but Fu discloses determining the screen text information based on the object image comprises: performing image processing on the object image to obtain third text information, (using OCR to perform text recognition on the caption, page 7); determining at least one piece of fourth text information based on the third text information, wherein semantics of the fourth text information is the same as, opposite to, or similar to semantics of the third text information (the words from the caption are filtered to produce key words, pages 7, 10; since these are filtered words, they have the same semantics as the corresponding words recognized from the text); and determining the third text information and the at least one piece of fourth text information as the screen text information (the recognized words of the caption selected as keywords for filtering are used as screen text information for recognition probabilities, page 10).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the invention of Corcoran such that determining the screen text information based on the object image comprises: performing image processing on the object image to obtain third text information,; determining at least one piece of fourth text information based on the third text information, wherein semantics of the fourth text information is the same as, opposite to, or similar to semantics of the third text information; and determining the third text information and the at least one piece of fourth text information as the screen text information for reasons similar to those for claim 1.
Corcoran, Fu, and Sarukkai do not specifically mention performing image description processing on the object image to obtain third text information, wherein the third text information describes the object in the object image.
Manjunath discloses performing image description processing on an object image to obtain text information, wherein the text information describes the object in the object image (extracting objects from the image frames and analyzing them to generate keywords indicating textual descriptors, [0051]).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the invention of Corcoran, Fu, and Sarukkai by performing image description processing on the object image to obtain third text information, wherein the third text information describes the object in the object image for reasons similar to those for claim 6.
Conclusion
Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a).
A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to Jesse Pullias whose telephone number is 571/270-5135. The examiner can normally be reached on M-F 8:00 AM - 4:30 PM. The examiner’s fax number is 571/270-6135.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner's supervisor, Andrew Flanders can be reached on 571/272-7516.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/Jesse S Pullias/
Primary Examiner, Art Unit 2655 08/07/26