DETAILED ACTION
This office action is in response to the above identified application filed on August 11, 2026. The application contains claims 1-21, wherein:
Claim 1 was previously cancelled
Claims 2-4, 7-10, 12-14, and 17-20 are amended
Claims 2-21 are pending
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Response to Arguments
Applicant's arguments and amendments filed on August 11, 2026, have been fully considered and the objections and rejections are updated accordingly.
Double Patenting
Applicant’s amendments to the claims do not overcome the double patenting rejections. As such, the nonstatutory double patenting rejections as set forth in the previous office action are maintained.
Because of the new matter and indefiniteness introduced with the new limitations, the content of the double patenting rejections is not duplicated in this office action. When all the issues are cleared, the double patenting rejections will be reevaluated against the amended claim language, and the content will be updated accordingly.
Claim Rejections - 35 USC § 101
Applicant’s amendments to the claims do not overcome the 35 U.S.C. 101 rejections.
In response to Applicant’s arguments on page 1 of Applicant’s Arguments/Remarks Made in an Amendment that the new limitations introduced with the amendments in claim 2 “improves the functioning of the computer itself”, the examiner disagrees.
The examiner notes, per MPEP 2106.05(a), "It is important to note that in order for a method claim to improve computer functionality, the broadest reasonable interpretation of the claim must be limited to computer implementation. That is, a claim whose entire scope can be performed mentally, cannot be said to improve computer technology. Synopsys, Inc. v. Mentor Graphics Corp., 839 F.3d 1138, 120 USPQ2d 1473 (Fed. Cir. 2016) (a method of translating a logic circuit into a hardware component description of a logic circuit was found to be ineligible because the method did not employ a computer and a skilled artisan could perform all the steps mentally). Similarly, a claimed process covering embodiments that can be performed on a computer, as well as embodiments that can be practiced verbally or with a telephone, cannot improve computer technology. See RecogniCorp, LLC v. Nintendo Co., 855 F.3d 1322, 1328, 122 USPQ2d 1377, 1381 (Fed. Cir. 2017) (process for encoding/decoding facial data using image codes assigned to particular facial features held ineligible because the process did not require a computer).” Because the entire scope of the claimed invention can be all performed in the human mind, the broadest reasonable interpretation of the claimed invention is not limited to computer implementation. The high-level recitation of generic computer components constitutes mere instructions to apply the abstract idea on a computer, which do not provide an inventive concept or significantly more.
Therefore, the 35 U.S.C. 101 rejections to claims 2-21 for being directed to an abstract idea are updated and maintained.
Claim Rejections - 35 USC § 103
Applicant argued about the new limitations introduced with the amendments, but as discussed in the 35 U.S.C. 112 rejections below, the new limitations recited in claim 2 have no support in the originally filed specification. Because of the lack of contextual information on how the new limitations fit in the claimed invention, the new limitation “selecting, using the control circuitry, a text representation corresponding to the one or more keywords based at least in part on one or more text representations accessed in a data structure of a user profile associated with the voice query” in claim 2 cannot be properly understood. As such, this new limitation is interpreted to mean using user profile information to help determine user audio queries and is addressed accordingly below.
Please refer to the updated 35 U.S.C. 103 rejections as set forth below for details.
Claim Objections
Claims 7, 9, 12, 17, and 19 are objected to because of the following informalities:
The limitation “the entity” recited in the following claims should be changed to “the first entity” to conform with the antecedent basis:
Claim 7, line 2
Claim 9, line 2
Claim 12, line 12
Claim 17, lines 1-2
Claim 19, line 1
Appropriate correction is required.
Claim Rejections - 35 USC § 112
The following is a quotation of the first paragraph of 35 U.S.C. 112(a):
(a) IN GENERAL.—The specification shall contain a written description of the invention, and of the manner and process of making and using it, in such full, clear, concise, and exact terms as to enable any person skilled in the art to which it pertains, or with which it is most nearly connected, to make and use the same, and shall set forth the best mode contemplated by the inventor or joint inventor of carrying out the invention.
The following is a quotation of the first paragraph of pre-AIA 35 U.S.C. 112:
The specification shall contain a written description of the invention, and of the manner and process of making and using it, in such full, clear, concise, and exact terms as to enable any person skilled in the art to which it pertains, or with which it is most nearly connected, to make and use the same, and shall set forth the best mode contemplated by the inventor of carrying out his invention.
Claims 2-11 are rejected under 35 U.S.C. 112(a) or 35 U.S.C. 112 (pre-AIA ), first paragraph, as failing to comply with the written description requirement. The claim(s) contains subject matter which was not described in the specification in such a way as to reasonably convey to one skilled in the relevant art that the inventor or a joint inventor, or for pre-AIA the inventor(s), at the time the application was filed, had possession of the claimed invention.
Claim 2 recites the limitation “selecting, using the control circuitry, a text representation corresponding to the one or more keywords based at least in part on one or more text representations accessed in a data structure of a user profile associated with the voice query” in lines 5-7. This limitation has no support in the specification as originally filed. Applicant pointed to paragraph [0019] in support of the amendment, but that paragraph discloses neither “a data structure of a user profile” nor the “selecting …”. Therefore, claim 2 is rejected under 35 U.S.C. 112(a).
Dependent claims 3-11 are also rejected for inheriting the deficiency from their corresponding independent claim 2, respectively.
The following is a quotation of 35 U.S.C. 112(b):
(b) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention.
The following is a quotation of 35 U.S.C. 112 (pre-AIA ), second paragraph:
The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the applicant regards as his invention.
Claims 2-21 are rejected under 35 U.S.C. 112(b) or 35 U.S.C. 112 (pre-AIA ), second paragraph, as being indefinite for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor, or for pre-AIA the applicant regards as the invention.
Claim 2 recites the limitation “selecting, using the control circuitry, a text representation corresponding to the one or more keywords based at least in part on one or more text representations accessed in a data structure of a user profile associated with the voice query” in lines 5-7. As discussed above, this limitation has no support in the originally filed specification. Lacking contextual information, this limitation cannot be properly understood. Therefore, claim 2 is indefinite and rejected under 35 U.S.C. 112(b).
Claim 12 recite “the text representation” as a claim limitation in lines 8-9. There is insufficient antecedent basis for this limitation in the claim. Therefore, claim 12 is indefinite and rejected under 35 U.S.C. 112(b).
Dependent claims 3-11 are also rejected for inheriting the deficiency from their corresponding independent claim 2, respectively.
Dependent claims 13-21 are also rejected for inheriting the deficiency from their corresponding independent claim 12, respectively.
Claim Rejections - 35 USC § 101
35 U.S.C. 101 reads as follows:
Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title.
Claims 2-21 are rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more.
The 2019 PEG guidance for subject matter eligibility is applied in the following analyses:
At Step 1
The inventions of claims 2-21 are directed to the statutory categories of a process (claims 2-11) and machine (claims 12-21). Thus, the claimed invention is directed to statutory subject matter.
At Step 2A, Prong One
The claimed invention is directed to mental processes without significantly more. Claims 2 and 12 recite abstract ideas in the following limitations:
“extracting … one or more keywords based at least in part on the voice query” recites a mental process as an evaluation or judgement of the important (i.e. key) words in audio. One listening to speech or audio can mentally evaluate that certain words heard are “important”, consistent with the specification at [0052].
“selecting, …, a text representation corresponding to the one or more keywords based at least in part on one or more text representations accessed in a data structure of a user profile associated with the voice query” recites a mental process as one may listen to speech or audio and mentally evaluate that a text representation associated with a user profile is a good match for a voice query. The high-level recitation of “a data structure”, which has no support in the specification as filed, constitutes mere instructions to implement an abstract idea on a computer or use a computer as a tool to perform an abstract idea, see MPEP 2106.05(f).
“generating … a text query based at least in part on the text representation” recites a mental process as one can mentally form a search query text based on a text representation, consistent with the specification at [0055].
“identifying a first entity of a plurality of entities based at least in part on a similarity of the text query to metadata stored for the first entity, wherein the metadata comprises one or more alternate text representations of an identifier associated with the first entity” recites a mental process as an evaluation (i.e. comparison) of the text query to stored alternate text representations of the entity based on pronunciation. Consistent with the specification at [0066], one can mentally compare text strings and determining a match.
At Step 2A, Prong Two
This judicial exception is not integrated into a practical application because the claims recite the additional elements of:
“an audio interface”, “control circuitry”, and “a memory” (claim 12) constitute a high-level recitation of a generic computer components and represent mere instructions to apply on a computer, see MPEP 2106.05(f).
“receiving a voice query” constitutes preliminary data gathering, see MPEP 2106.05(g).
“retrieving a content item associated with the first entity” constitutes preliminary data gathering, see MPEP 2106.05(g) or as mere instruction to ‘apply it’ under MPEP 2106.05(f).
“store in the memory the voice query” (claim 12) may be characterized as insignificant extra-solution activity, particularly post-solution activity, see MPEP 2106.05(g).
Even when viewed in combination, these additional elements do not integrate the recited judicial exception into a practical application and the claim is directed to the judicial exception.
At Step 2B
Claims 2 and 12 do not include additional elements that are sufficient to amount to significantly more than the judicial exception because as discussed above the additional elements constitute a high-level recitation of a generic computer components which represent mere instructions to apply on a computer, preliminary data gathering, and insignificant extra-solution activity, particularly post-solution activity. As identified by courts retrieving, receiving, and storing data are well-understood, routine, and conventional activities, see MPEP 2106.05(d). [Receiving or transmitting data over a network, e.g., using the Internet to gather data, Symantec, 838 F.3d at 1321, 120 USPQ2d at 1362 (utilizing an intermediary computer to forward information); Storing and retrieving information in memory, Versata Dev. Group, Inc. v. SAP Am., Inc., 793 F.3d 1306, 1334, 115 USPQ2d 1681, 1701 (Fed. Cir. 2015); OIP Techs., 788 F.3d at 1363, 115 USPQ2d at 1092-93;].
Even when considered in combination, these additional elements do not provide an inventive concept or significantly more.
Therefore, claims 2 and 12 are rejected under 35 USC 101 as being directed to an abstract idea without significantly more.
Dependent claims 3-11 and 13-21 each recite abstract ideas elaborating on the further details of the “identifying” in the independent claims 2 and 12 that are still mentally performable.
Therefore, dependent claims 3-11 and 13-21 are also rejected under 35 USC 101 as being directed to an abstract idea without significantly more.
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claims 2, 3, 5, 7-10, 12, 13, 15, and 17-20 are rejected under 35 U.S.C. 103 as being unpatentable over Jang (US 20140359523 A1), in view of Ramos et al. (US 11157696 B1), and in further view of Olstad et al. (US 20130132374 A1).
With regard to claim 2,
Jang teaches
a method for responding to voice queries (Fig. 11; [0165]-[0169]; Fig. 12; [0170]-[0173]), the method comprising:
receiving a voice query received at an audio interface ([0166]; Fig. 1, microphone 122; [0047]: receive a voice query through an audio input component such as a microphone, wherein the microphone corresponds to “an audio interface”);
extracting, using control circuitry (Fig. 1; [0073]-[0076]: controller 180 corresponds to “control circuity”), one or more keywords based at least in part on the voice query ([0170]: identify a query term of the voice query, wherein a query term corresponds to “one or more keywords based at least in part on the voice query”);
generating, using the control circuitry, a text query based at least in part on the text representation (Fig. 11; [0167]-[0168]; Fig. 12; [0170]: convert the voice query to a text query, which identifies query terms from the voice query, determines pronunciation information for each query term, and converts each query term into a typical text query term using a voice query term database that links a range of pronunciation of terms to a typical query term. As a result, a text query is generated comprising the query terms and their pronunciations);
Jang does not explicitly teach
selecting, using the control circuitry, a text representation corresponding to the one or more keywords based at least in part on one or more text representations accessed in a data structure of a user profile associated with the voice query;
identifying a first entity of a plurality of entities based at least in part on a similarity of the text query to metadata stored for the first entity, wherein the metadata comprises one or more alternate text representations of an identifier associated with the first entity; and
retrieving a content item associated with the first entity.
Ramos teaches
selecting, using the control circuitry, a text representation corresponding to the one or more keywords based at least in part on one or more text representations accessed in a data structure of a user profile associated with the voice query (due to the indefiniteness discussed above, this limitation is interpreted to mean using user profile information to help determine user audio queries. This is taught by claim 2);
identifying a first entity of a plurality of entities based at least in part on a similarity of the text query to metadata stored for the first entity, wherein the metadata comprises a pronunciation tag comprising a phonetic spelling for the entity (Fig. 1, 138-144; Col. 3, lines 61-67; Col. 4, lines 1-56: perform entity resolution based on the tagged portion of text data and the portion of audio data corresponding to the tagged portion of text by comparing the portion of audio data against audio data representing entities known to the system, wherein performing entity resolution corresponds to “identifying a first entity”, the tagged portion of text data corresponds to “the text query”, and audio data representing entities known to the system corresponds to “a pronunciation tag” comprised in the “metadata stored for the entity”. Fig. 6; Col. 16, lines 42-67; Col. 17, lines 1-18: audio data comprises phonetic representation of text data, wherein the phonetic representation reads on "phonetic spelling"); and
retrieving a content item associated with the first entity (Fig. 1, 146; Col. 4, lines 57-62; Col. 2, lines 10-18: use the resolved entity to perform downstream processes. For example, for the user input of "Alexa, play Adele music," a system may output music sung by Adele, wherein output indicates “retrieving”, and music sung by Adele corresponds to “a content item associated with the first entity” Adele).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified Jang to incorporate the teachings of Ramos to identify an entity based at least in part on the text query and metadata for the entity, wherein the metadata comprises one or more alternate text representations of an identifier associated with the entity and retrieve a content item associated with the entity. Doing so would improve text-based entity resolution by providing a language agnostic phonetic searching as part of entity resolution when text-based entity resolution may be unsuccessful or successful to a degree below a requisite threshold confidence as taught by Ramos (Col. 2, lines 48-67).
Jang and Ramos do not teach
identifying a first entity of a plurality of entities based at least in part on a similarity of the text query to metadata stored for the first entity, wherein the metadata comprises one or more alternate text representations of an identifier associated with the first entity;
Olstad teaches
identifying a first entity of a plurality of entities based at least in part on a similarity of the text query to metadata stored for the first entity, wherein the metadata comprises one or more alternate text representations of an identifier associated with the first entity (Claim 20; [0039]: identify videos in the set of videos based on metadata associated with the videos, wherein the metadata includes phonetic transcription extracted from the audio track, and the phoneme sequences included in a phonetic transcription of the audio track are matched with a phonetic representation of the query to find locations inside the audio track with the best phonetic similarity. “query terms” indicates “text query” and phonetic transcription metadata is “one or more alternate text representations” of a speech-to-text transcription associated with the identified videos, i.e., “an identifier”);
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified Jang and Ramos to incorporate the teachings of Olstad to identify an entity based at least in part on the text query and metadata for the entity, wherein the metadata comprises one or more alternate text representations of an identifier associated with the entity. Doing so would find locations inside the audio track with the best phonetic similarity to a user query, improve search precision, and perform less analysis including metadata generation as taught by Olstad ([0039]).
With regard to claim 3,
As discussed in claim 2, Jang and Ramos and Olstad teach all the limitations therein.
Olstad further teaches
the method of claim 2, wherein the one or more alternate text representations comprises a phonetic representation of the identifier associated with the first entity (Claim 20; [0039]: the phonetic transcription included in the metadata that is extracted from the audio track of the identified videos corresponds to “a phonetic representation”).
With regard to claim 5,
As discussed in claim 2, Jang and Ramos and Olstad teach all the limitations therein.
Olstad further teaches
the method of claim 2, wherein the one or more alternate text representations comprises a text string generated based at least in part on a previous speech-to-text conversion. ([0039]: phonetic transcription being an alternative to speech-to-text transcription of the audio track in the video indicates generating a text string by either speech-to-text conversion or phonetic transcription, which takes place before the metadata is used for search, i.e., “a previous”).
With regard to claim 7,
As discussed in claim 2, Jang and Ramos and Olstad teach all the limitations therein.
Olstad further teaches
the method of claim 2, wherein the identifier associated with the first entity identifies information related to the entity (Claim 20; [0039]: the phonetic transcription included in the metadata of the identified videos identifies the audio track in the videos, i.e., “identifies information related to” the videos, i.e., “the first entity”).
With regard to claim 8,
As discussed in claim 2, Jang and Ramos and Olstad teach all the limitations therein.
Ramos further teaches
the method of claim 2, wherein the identifying the first entity is based at least in part on user profile information. (Col. 20, lines 13-19; Col. 7, lines 61-66: the phonetic entity resolution component 802 may consider user preferences, wherein user preferences read on “user profile information”).
With regard to claim 9,
As discussed in claim 2, Jang and Ramos and Olstad teach all the limitations therein.
Ramos further teaches
the method of claim 2, wherein the identifying the first entity is based at least in part on popularity information associated with the entity. (Col. 20, lines 13-19: the phonetic entity resolution component 802 may consider popularity of known entities).
With regard to claim 10,
As discussed in claim 2, Jang and Ramos and Olstad teach all the limitations therein.
Olstad further teaches
the method of claim 2, further comprising:
identifying a second entity based at least in part on the text query and metadata for the second entity (Claim 20; [0039]; [0003]: querying videos by matching a phonetic representation of the query to phonetic transcription of the audio track inside the videos indicates identifying plural videos that include a first video, a second video, and potentially more, wherein each video corresponds to an “entity”), and
Ramos further teaches
determining a first score for the first entity based at least in part on a comparison of the text query to the metadata associated with the first entity, and determining a second score for the second entity based at least in part on a comparison of the text query to metadata associated with the second entity, wherein the content item associated with the first entity is retrieved by selecting a maximum score of the first score and the second score (Fig. 8; Col. 19, lines 27-41: perform phonetic matching of the audio data (representing the entity to be resolved) to audio data stored in the entity storage (608/706) and associate with each known entity a confidence value representing the confidence that the known entity corresponds to the entity in the user input, wherein the confidence value corresponds to a “score” and the phonetic matching corresponds to a “comparison”. Col. 20, lines 20-41: the N-best list may include a maximum number of top scoring known entities, wherein selecting the maximum number of top scoring known entities indicates not only every entity is scored but also only the ones with "a maximum score" gets selected).
With regard to claim 12,
Jang teaches
a system for responding to voice queries (Fig. 11; [0165]-[0169]; Fig. 12; [0170]-[0173]), the system comprising:
a memory (Fig. 1: memory 160); and
control circuitry (Fig. 1; [0073]-[0076]: controller 180 corresponds to “control circuity”) configured to:
receive a voice query received at an audio interface ([0166]; Fig. 1, microphone 122; [0047]: receive a voice query through an audio input component such as a microphone, wherein the microphone corresponds to “an audio interface”);
store in the memory the voice query (since the memory is the place for data a computer needs to access quickly for processing, receiving a voice query as discussed above in [0166]; Fig. 1, microphone 122; [0047] and processing the voice query to extract keywords as discussed below in [0170] both inherently teach storing the received voice query in the memory);
extract one or more keywords based at least in part on the voice query ([0170]: identify a query term of the voice query, wherein a query term corresponds to “one or more keywords based at least in part on the voice query”);
generate a text query based at least in part on the text representation (Fig. 11; [0167]-[0168]; Fig. 12; [0170]: convert the voice query to a text query, which identifies query terms from the voice query, determines pronunciation information for each query term, and converts each query term into a typical text query term using a voice query term database that links a range of pronunciation of terms to a typical query term. As a result, a text query is generated comprising the query terms and their pronunciations);
Jang does not explicitly teach
identify a first entity of a plurality of entities based at least in part on a similarity of the text query to metadata stored for the entity, wherein the metadata comprises one or more alternate text representations of an identifier associated with the first entity; and
retrieve a content item associated with the first entity.
Ramos teaches
identify a first entity of a plurality of entities based at least in part on a similarity of the text query to metadata stored for the entity, wherein the metadata comprises a pronunciation tag comprising a phonetic spelling for the entity (Fig. 1, 138-144; Col. 3, lines 61-67; Col. 4, lines 1-56: perform entity resolution based on the tagged portion of text data and the portion of audio data corresponding to the tagged portion of text by comparing the portion of audio data against audio data representing entities known to the system, wherein performing entity resolution corresponds to “identifying a first entity”, the tagged portion of text data corresponds to “the text query”, and audio data representing entities known to the system corresponds to “a pronunciation tag” comprised in the “metadata stored for the entity”. Fig. 6; Col. 16, lines 42-67; Col. 17, lines 1-18: audio data comprises phonetic representation of text data, wherein the phonetic representation reads on "phonetic spelling"); and
retrieve a content item associated with the first entity (Fig. 1, 146; Col. 4, lines 57-62; Col. 2, lines 10-18: use the resolved entity to perform downstream processes. For example, for the user input of "Alexa, play Adele music," a system may output music sung by Adele, wherein output indicates “retrieving”, and music sung by Adele corresponds to “a content item associated with the first entity” Adele).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified Jang to incorporate the teachings of Ramos to identify an entity based at least in part on the text query and metadata for the entity, wherein the metadata comprises one or more alternate text representations of an identifier associated with the entity and retrieve a content item associated with the entity. Doing so would improve text-based entity resolution by providing a language agnostic phonetic searching as part of entity resolution when text-based entity resolution may be unsuccessful or successful to a degree below a requisite threshold confidence as taught by Ramos (Col. 2, lines 48-67).
Jang and Ramos do not teach
identify a first entity of a plurality of entities based at least in part on a similarity of the text query to metadata stored for the entity, wherein the metadata comprises one or more alternate text representations of an identifier associated with the first entity;
Olstad teaches
identify a first entity of a plurality of entities based at least in part on a similarity of the text query to metadata stored for the entity, wherein the metadata comprises one or more alternate text representations of an identifier associated with the first entity (Claim 20; [0039]: identify videos in the set of videos based on metadata associated with the videos, wherein the metadata includes phonetic transcription extracted from the audio track, and the phoneme sequences included in a phonetic transcription of the audio track are matched with a phonetic representation of the query to find locations inside the audio track with the best phonetic similarity. “query terms” indicates “text query” and phonetic transcription metadata is “one or more alternate text representations” of a speech-to-text transcription associated with the identified videos, i.e., “an identifier”);
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified Jang and Ramos to incorporate the teachings of Olstad to identify an entity based at least in part on the text query and metadata for the entity, wherein the metadata comprises one or more alternate text representations of an identifier associated with the entity. Doing so would find locations inside the audio track with the best phonetic similarity to a user query, improve search precision, and perform less analysis including metadata generation as taught by Olstad ([0039]).
With regard to claim 13,
As discussed in claim 12, Jang and Ramos and Olstad teach all the limitations therein.
Olstad further teaches
the system of claim 12, wherein the one or more alternate text representations comprises a phonetic representation of the identifier associated with the first entity (Claim 20; [0039]: the phonetic transcription included in the metadata that is extracted from the audio track of the identified videos corresponds to “a phonetic representation”).
With regard to claim 15,
As discussed in claim 12, Jang and Ramos and Olstad teach all the limitations therein.
Olstad further teaches
the system of claim 12, wherein the one or more alternate text representations comprises a text string generated based at least in part on a previous speech-to-text conversion. ([0039]: phonetic transcription being an alternative to speech-to-text transcription of the audio track in the video indicates generating a text string by either speech-to-text conversion or phonetic transcription, which takes place before the metadata is used for search, i.e., “a previous”).
With regard to claim 17,
As discussed in claim 12, Jang and Ramos and Olstad teach all the limitations therein.
Olstad further teaches
the system of claim 12, wherein the identifier associated with the entity identifies information related to the first entity (Claim 20; [0039]: the phonetic transcription included in the metadata of the identified videos identifies the audio track in the videos, i.e., “identifies information related to” the videos, i.e., “the first entity”).
With regard to claim 18,
As discussed in claim 12, Jang and Ramos and Olstad teach all the limitations therein.
Ramos further teaches
the system of claim 12, wherein the identifying the first entity is based at least in part on user profile information. (Col. 20, lines 13-19; Col. 7, lines 61-66: the phonetic entity resolution component 802 may consider user preferences, wherein user preferences read on “user profile information”).
With regard to claim 19,
As discussed in claim 12, Jang and Ramos and Olstad teach all the limitations therein.
Ramos further teaches
the system of claim 12, wherein the identifying the entity is based at least in part on popularity information associated with the first entity. (Col. 20, lines 13-19: the phonetic entity resolution component 802 may consider popularity of known entities).
With regard to claim 20,
As discussed in claim 12, Jang and Ramos and Olstad teach all the limitations therein.
Olstad further teaches
the system of claim 12, the system is configured to:
identify a second entity based at least in part on the text query and metadata for the second entity (Claim 20; [0039]; [0003]: querying videos by matching a phonetic representation of the query to phonetic transcription of the audio track inside the videos indicates identifying plural videos that include a first video, a second video, and potentially more, wherein each video corresponds to an “entity”), and
Ramos further teaches
determine a first score for the first entity based at least in part on a comparison of the text query to the metadata associated with the first entity, and determine a second score for the second entity based at least in part on a comparison of the text query to metadata associated with the second entity, wherein the content item associated with the first entity is retrieved by selecting a maximum score of the first score and the second score (Fig. 8; Col. 19, lines 27-41: perform phonetic matching of the audio data (representing the entity to be resolved) to audio data stored in the entity storage (608/706) and associate with each known entity a confidence value representing the confidence that the known entity corresponds to the entity in the user input, wherein the confidence value corresponds to a “score” and the phonetic matching corresponds to a “comparison”. Col. 20, lines 20-41: the N-best list may include a maximum number of top scoring known entities, wherein selecting the maximum number of top scoring known entities indicates not only every entity is scored but also only the ones with "a maximum score" gets selected).
Claims 4 and 14 are rejected under 35 U.S.C. 103 as being unpatentable over Jang (US 20140359523 A1), in view of Ramos et al. (US 11157696 B1), and in further view of Olstad et al. (US 20130132374 A1) and YAO et al. (US 20110307432 A1).
With regard to claim 4,
As discussed in claim 2, Jang and Ramos and Olstad teach all the limitations therein.
Jang and Ramos and Olstad do not teach
the method of claim 2, wherein the one or more alternate text representations comprises an alternate spelling of the identifier associated with the first entity.
YAO teaches
the method of claim 2, wherein the one or more alternate text representations comprises an alternate spelling of the identifier associated with the first entity ([0055]: indexed metadata for web pages includes entity name equivalents data, variations of an entity's name, and entity name misspellings, all of which correspond to "an alternate spelling" associated with the web pages).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified Jang and Ramos and Olstad to incorporate the teachings of YAO to make the one or more alternate text representations comprise an alternate spelling of the identifier associated with the entity. Doing so would provide improved search result relevance for name search queries as taught by YAO ([0005]).
With regard to claim 14,
As discussed in claim 12, Jang and Ramos and Olstad teach all the limitations therein.
Jang and Ramos and Olstad do not teach
the system of claim 12, wherein the one or more alternate text representations comprises an alternate spelling of the identifier associated with the first entity.
YAO teaches
the system of claim 12, wherein the one or more alternate text representations comprises an alternate spelling of the identifier associated with the first entity ([0055]: indexed metadata for web pages includes entity name equivalents data, variations of an entity's name, and entity name misspellings, all of which correspond to "an alternate spelling" associated with the web pages).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified Jang and Ramos and Olstad to incorporate the teachings of YAO to make the one or more alternate text representations comprise an alternate spelling of the identifier associated with the entity. Doing so would provide improved search result relevance for name search queries as taught by YAO ([0005]).
Claims 6 and 16 are rejected under 35 U.S.C. 103 as being unpatentable over Jang (US 20140359523 A1), in view of Ramos et al. (US 11157696 B1), and in further view of Olstad et al. (US 20130132374 A1) and Pore et al. (US 20190295527 A1).
With regard to claim 6,
As discussed in claim 2, Jang and Ramos and Olstad teach all the limitations therein.
Jang and Ramos and Olstad do not teach
the method of claim 2, wherein the metadata is generated by a text-to-speech module using a pronunciation setting for generating speech content; and
converting the generated speech content into text using a speech-to-text module.
Pore teaches
the method of claim 2, wherein the metadata is generated by a text-to-speech module using a pronunciation setting for generating speech content; and converting the generated speech content into text using a speech-to-text module (Abstract; [0035]-[0036]: convert a text message containing at least one phonemic spelling of a word into speech by running a text-to-speech application programming interface (API) with the text message as input. The converted speech may be input to a speech-to-text API and the speech-to-text API executed to convert the speech to text. Users of English language in different geographic locations may have different accents or pronunciations, and therefore, may pronounce or voice an English word based on phonemes particular to the geographic location. A speech-to-text API that can recognize the particular location's accents or pronunciation of words may provide for a more accurate conversion of speech into text).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified Jang and Ramos and Olstad to incorporate the teachings of Pore to generate the metadata by a text-to-speech module using a pronunciation setting for generating speech content, and convert the generated speech content into text using a speech-to-text module. Doing so would recognize a particular location's accents of a user and provide more accurate conversion of the text into English speech as taught by Pore ([0035]).
With regard to claim 16,
As discussed in claim 12, Jang and Ramos and Olstad teach all the limitations therein.
Jang and Ramos and Olstad do not teach
the system of claim 12, wherein the metadata is generated by a text-to-speech module using a pronunciation setting for generating speech content; and the system is configured to:
converting the generated speech content into text using a speech-to-text module.
Pore teaches
the system of claim 12, wherein the metadata is generated by a text-to-speech module using a pronunciation setting for generating speech content; and the system is configured to: converting the generated speech content into text using a speech-to-text module (Abstract; [0035]-[0036]: convert a text message containing at least one phonemic spelling of a word into speech by running a text-to-speech application programming interface (API) with the text message as input. The converted speech may be input to a speech-to-text API and the speech-to-text API executed to convert the speech to text. Users of English language in different geographic locations may have different accents or pronunciations, and therefore, may pronounce or voice an English word based on phonemes particular to the geographic location. A speech-to-text API that can recognize the particular location's accents or pronunciation of words may provide for a more accurate conversion of speech into text).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified Jang and Ramos and Olstad to incorporate the teachings of Pore to generate the metadata by a text-to-speech module using a pronunciation setting for generating speech content, and convert the generated speech content into text using a speech-to-text module. Doing so would recognize a particular location's accents of a user and provide more accurate conversion of the text into English speech as taught by Pore ([0035]).
Claims 11 and 21 are rejected under 35 U.S.C. 103 as being unpatentable over Jang (US 20140359523 A1), in view of Ramos et al. (US 11157696 B1), and in further view of Olstad et al. (US 20130132374 A1) and Davallou (US 20060074892 A1).
With regard to claim 11,
As discussed in claim 2, Jang and Ramos and Olstad teach all the limitations therein.
Jang and Ramos and Olstad do not teach
the method of claim 2, wherein the text query is a first text query, and further comprising:
generating a plurality of text queries, wherein the plurality of text queries comprises the first text query, and wherein each text query of the plurality of text queries is generated based at least in part on a respective pronunciation setting of a speech-to-text module.
Davallou teaches
the method of claim 2, wherein the text query is a first text query, and further comprising:
generating a plurality of text queries, wherein the plurality of text queries comprises the first text query, and wherein each text query of the plurality of text queries is generated based at least in part on a respective pronunciation setting of a speech-to-text module ([0064]: generate the search string through a speech-to-text software application from the dictation of a user interacting with a database, such as through a telephone system or microphone. [0079]; [0069]: search for similar sounding words within the phonetic database for each word and generate a number of different combinations based on the approximate pronunciation of the text entered and the phonetically equivalent formulas, the result of which would be “generating a plurality of text queries” based on a pronunciation setting of a speech-to-text module).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified Jang and Ramos and Olstad to incorporate the teachings of Davallou to generate a plurality of text queries, wherein the plurality of text queries comprises the first text query, and wherein each text query of the plurality of text queries is generated based at least in part on a respective pronunciation setting of a speech-to-text module. Doing so would eventually find the correct query even though the original entry involved a different spelling, without the user having to respell or retype the entry as taught by Davallou ([0069]).
With regard to claim 21,
As discussed in claim 12, Jang and Ramos and Olstad teach all the limitations therein.
Jang and Ramos and Olstad do not teach
the system of claim 12, wherein the text query is a first text query, and further comprising:
generating a plurality of text queries, wherein the plurality of text queries comprises the first text query, and wherein each text query of the plurality of text queries is generated based at least in part on a respective pronunciation setting of a speech-to-text module.
Davallou teaches
the system of claim 12, wherein the text query is a first text query, and further comprising:
generating a plurality of text queries, wherein the plurality of text queries comprises the first text query, and wherein each text query of the plurality of text queries is generated based at least in part on a respective pronunciation setting of a speech-to-text module ([0064]: generate the search string through a speech-to-text software application from the dictation of a user interacting with a database, such as through a telephone system or microphone. [0079]; [0069]: search for similar sounding words within the phonetic database for each word and generate a number of different combinations based on the approximate pronunciation of the text entered and the phonetically equivalent formulas, the result of which would be “generating a plurality of text queries” based on a pronunciation setting of a speech-to-text module).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified Jang and Ramos and Olstad to incorporate the teachings of Davallou to generate a plurality of text queries, wherein the plurality of text queries comprises the first text query, and wherein each text query of the plurality of text queries is generated based at least in part on a respective pronunciation setting of a speech-to-text module. Doing so would eventually find the correct query even though the original entry involved a different spelling, without the user having to respell or retype the entry as taught by Davallou ([0069]).
Conclusion
Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a).
A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any extension fee pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the date of this final action.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to XIAOQIN HU whose telephone number is (571)272-1792. The examiner can normally be reached on Monday-Friday 7:00am-3:30pm.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Charles Rones can be reached on (571) 272-4085. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of an application may be obtained from the Patent Application Information Retrieval (PAIR) system. Status information for published applications may be obtained from either Private PAIR or Public PAIR. Status information for unpublished applications is available through Private PAIR only. For more information about the PAIR system, see http://pair-direct.uspto.gov. Should you have questions on access to the Private PAIR system, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative or access to the automated information system, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/XIAOQIN HU/Examiner, Art Unit 2168
/CHARLES RONES/Supervisory Patent Examiner, Art Unit 2168