Prosecution Insights
Last updated: October 02, 2026
Application No. 19/211,748

SYSTEMS AND METHODS FOR MANAGING VOICE QUERIES USING PRONUNCIATION INFORMATION

Final Rejection §101§103§112
Filed
May 19, 2025
Priority
Jul 31, 2019 — continuation of 12/332,937
Examiner
HU, XIAOQIN
Art Unit
2168
Tech Center
2100 — Computer Architecture & Software
Assignee
Adeia Technologies Inc.
OA Round
2 (Final)
62%
Grant Probability
Moderate
3-4
OA Rounds
1y 6m
Est. Remaining
99%
With Interview

Examiner Intelligence

Grants 62% of resolved cases
62%
Career Allowance Rate
120 granted / 195 resolved
+6.5% vs TC avg
Strong +56% interview lift
Without
With
+56.2%
Interview Lift
resolved cases with interview
Typical timeline
2y 10m
Avg Prosecution
20 currently pending
Career history
222
Total Applications
across all art units

Statute-Specific Performance

§101
17.1%
-22.9% vs TC avg
§103
40.9%
+0.9% vs TC avg
§102
10.6%
-29.4% vs TC avg
§112
28.7%
-11.3% vs TC avg
Black line = Tech Center average estimate • Based on career data from 195 resolved cases

Office Action

§101 §103 §112
DETAILED ACTION This office action is in response to the above identified application filed on August 11, 2026. The application contains claims 1-21, wherein: Claim 1 was previously cancelled Claims 2-4, 7-10, 12-14, and 17-20 are amended Claims 2-21 are pending Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Response to Arguments Applicant's arguments and amendments filed on August 11, 2026, have been fully considered and the objections and rejections are updated accordingly. Double Patenting Applicant’s amendments to the claims do not overcome the double patenting rejections. As such, the nonstatutory double patenting rejections as set forth in the previous office action are maintained. Because of the new matter and indefiniteness introduced with the new limitations, the content of the double patenting rejections is not duplicated in this office action. When all the issues are cleared, the double patenting rejections will be reevaluated against the amended claim language, and the content will be updated accordingly. Claim Rejections - 35 USC § 101 Applicant’s amendments to the claims do not overcome the 35 U.S.C. 101 rejections. In response to Applicant’s arguments on page 1 of Applicant’s Arguments/Remarks Made in an Amendment that the new limitations introduced with the amendments in claim 2 “improves the functioning of the computer itself”, the examiner disagrees. The examiner notes, per MPEP 2106.05(a), "It is important to note that in order for a method claim to improve computer functionality, the broadest reasonable interpretation of the claim must be limited to computer implementation. That is, a claim whose entire scope can be performed mentally, cannot be said to improve computer technology. Synopsys, Inc. v. Mentor Graphics Corp., 839 F.3d 1138, 120 USPQ2d 1473 (Fed. Cir. 2016) (a method of translating a logic circuit into a hardware component description of a logic circuit was found to be ineligible because the method did not employ a computer and a skilled artisan could perform all the steps mentally). Similarly, a claimed process covering embodiments that can be performed on a computer, as well as embodiments that can be practiced verbally or with a telephone, cannot improve computer technology. See RecogniCorp, LLC v. Nintendo Co., 855 F.3d 1322, 1328, 122 USPQ2d 1377, 1381 (Fed. Cir. 2017) (process for encoding/decoding facial data using image codes assigned to particular facial features held ineligible because the process did not require a computer).” Because the entire scope of the claimed invention can be all performed in the human mind, the broadest reasonable interpretation of the claimed invention is not limited to computer implementation. The high-level recitation of generic computer components constitutes mere instructions to apply the abstract idea on a computer, which do not provide an inventive concept or significantly more. Therefore, the 35 U.S.C. 101 rejections to claims 2-21 for being directed to an abstract idea are updated and maintained. Claim Rejections - 35 USC § 103 Applicant argued about the new limitations introduced with the amendments, but as discussed in the 35 U.S.C. 112 rejections below, the new limitations recited in claim 2 have no support in the originally filed specification. Because of the lack of contextual information on how the new limitations fit in the claimed invention, the new limitation “selecting, using the control circuitry, a text representation corresponding to the one or more keywords based at least in part on one or more text representations accessed in a data structure of a user profile associated with the voice query” in claim 2 cannot be properly understood. As such, this new limitation is interpreted to mean using user profile information to help determine user audio queries and is addressed accordingly below. Please refer to the updated 35 U.S.C. 103 rejections as set forth below for details. Claim Objections Claims 7, 9, 12, 17, and 19 are objected to because of the following informalities: The limitation “the entity” recited in the following claims should be changed to “the first entity” to conform with the antecedent basis: Claim 7, line 2 Claim 9, line 2 Claim 12, line 12 Claim 17, lines 1-2 Claim 19, line 1 Appropriate correction is required. Claim Rejections - 35 USC § 112 The following is a quotation of the first paragraph of 35 U.S.C. 112(a): (a) IN GENERAL.—The specification shall contain a written description of the invention, and of the manner and process of making and using it, in such full, clear, concise, and exact terms as to enable any person skilled in the art to which it pertains, or with which it is most nearly connected, to make and use the same, and shall set forth the best mode contemplated by the inventor or joint inventor of carrying out the invention. The following is a quotation of the first paragraph of pre-AIA 35 U.S.C. 112: The specification shall contain a written description of the invention, and of the manner and process of making and using it, in such full, clear, concise, and exact terms as to enable any person skilled in the art to which it pertains, or with which it is most nearly connected, to make and use the same, and shall set forth the best mode contemplated by the inventor of carrying out his invention. Claims 2-11 are rejected under 35 U.S.C. 112(a) or 35 U.S.C. 112 (pre-AIA ), first paragraph, as failing to comply with the written description requirement. The claim(s) contains subject matter which was not described in the specification in such a way as to reasonably convey to one skilled in the relevant art that the inventor or a joint inventor, or for pre-AIA the inventor(s), at the time the application was filed, had possession of the claimed invention. Claim 2 recites the limitation “selecting, using the control circuitry, a text representation corresponding to the one or more keywords based at least in part on one or more text representations accessed in a data structure of a user profile associated with the voice query” in lines 5-7. This limitation has no support in the specification as originally filed. Applicant pointed to paragraph [0019] in support of the amendment, but that paragraph discloses neither “a data structure of a user profile” nor the “selecting …”. Therefore, claim 2 is rejected under 35 U.S.C. 112(a). Dependent claims 3-11 are also rejected for inheriting the deficiency from their corresponding independent claim 2, respectively. The following is a quotation of 35 U.S.C. 112(b): (b) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention. The following is a quotation of 35 U.S.C. 112 (pre-AIA ), second paragraph: The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the applicant regards as his invention. Claims 2-21 are rejected under 35 U.S.C. 112(b) or 35 U.S.C. 112 (pre-AIA ), second paragraph, as being indefinite for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor, or for pre-AIA the applicant regards as the invention. Claim 2 recites the limitation “selecting, using the control circuitry, a text representation corresponding to the one or more keywords based at least in part on one or more text representations accessed in a data structure of a user profile associated with the voice query” in lines 5-7. As discussed above, this limitation has no support in the originally filed specification. Lacking contextual information, this limitation cannot be properly understood. Therefore, claim 2 is indefinite and rejected under 35 U.S.C. 112(b). Claim 12 recite “the text representation” as a claim limitation in lines 8-9. There is insufficient antecedent basis for this limitation in the claim. Therefore, claim 12 is indefinite and rejected under 35 U.S.C. 112(b). Dependent claims 3-11 are also rejected for inheriting the deficiency from their corresponding independent claim 2, respectively. Dependent claims 13-21 are also rejected for inheriting the deficiency from their corresponding independent claim 12, respectively. Claim Rejections - 35 USC § 101 35 U.S.C. 101 reads as follows: Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title. Claims 2-21 are rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more. The 2019 PEG guidance for subject matter eligibility is applied in the following analyses: At Step 1 The inventions of claims 2-21 are directed to the statutory categories of a process (claims 2-11) and machine (claims 12-21). Thus, the claimed invention is directed to statutory subject matter. At Step 2A, Prong One The claimed invention is directed to mental processes without significantly more. Claims 2 and 12 recite abstract ideas in the following limitations: “extracting … one or more keywords based at least in part on the voice query” recites a mental process as an evaluation or judgement of the important (i.e. key) words in audio. One listening to speech or audio can mentally evaluate that certain words heard are “important”, consistent with the specification at [0052]. “selecting, …, a text representation corresponding to the one or more keywords based at least in part on one or more text representations accessed in a data structure of a user profile associated with the voice query” recites a mental process as one may listen to speech or audio and mentally evaluate that a text representation associated with a user profile is a good match for a voice query. The high-level recitation of “a data structure”, which has no support in the specification as filed, constitutes mere instructions to implement an abstract idea on a computer or use a computer as a tool to perform an abstract idea, see MPEP 2106.05(f). “generating … a text query based at least in part on the text representation” recites a mental process as one can mentally form a search query text based on a text representation, consistent with the specification at [0055]. “identifying a first entity of a plurality of entities based at least in part on a similarity of the text query to metadata stored for the first entity, wherein the metadata comprises one or more alternate text representations of an identifier associated with the first entity” recites a mental process as an evaluation (i.e. comparison) of the text query to stored alternate text representations of the entity based on pronunciation. Consistent with the specification at [0066], one can mentally compare text strings and determining a match. At Step 2A, Prong Two This judicial exception is not integrated into a practical application because the claims recite the additional elements of: “an audio interface”, “control circuitry”, and “a memory” (claim 12) constitute a high-level recitation of a generic computer components and represent mere instructions to apply on a computer, see MPEP 2106.05(f). “receiving a voice query” constitutes preliminary data gathering, see MPEP 2106.05(g). “retrieving a content item associated with the first entity” constitutes preliminary data gathering, see MPEP 2106.05(g) or as mere instruction to ‘apply it’ under MPEP 2106.05(f). “store in the memory the voice query” (claim 12) may be characterized as insignificant extra-solution activity, particularly post-solution activity, see MPEP 2106.05(g). Even when viewed in combination, these additional elements do not integrate the recited judicial exception into a practical application and the claim is directed to the judicial exception. At Step 2B Claims 2 and 12 do not include additional elements that are sufficient to amount to significantly more than the judicial exception because as discussed above the additional elements constitute a high-level recitation of a generic computer components which represent mere instructions to apply on a computer, preliminary data gathering, and insignificant extra-solution activity, particularly post-solution activity. As identified by courts retrieving, receiving, and storing data are well-understood, routine, and conventional activities, see MPEP 2106.05(d). [Receiving or transmitting data over a network, e.g., using the Internet to gather data, Symantec, 838 F.3d at 1321, 120 USPQ2d at 1362 (utilizing an intermediary computer to forward information); Storing and retrieving information in memory, Versata Dev. Group, Inc. v. SAP Am., Inc., 793 F.3d 1306, 1334, 115 USPQ2d 1681, 1701 (Fed. Cir. 2015); OIP Techs., 788 F.3d at 1363, 115 USPQ2d at 1092-93;]. Even when considered in combination, these additional elements do not provide an inventive concept or significantly more. Therefore, claims 2 and 12 are rejected under 35 USC 101 as being directed to an abstract idea without significantly more. Dependent claims 3-11 and 13-21 each recite abstract ideas elaborating on the further details of the “identifying” in the independent claims 2 and 12 that are still mentally performable. Therefore, dependent claims 3-11 and 13-21 are also rejected under 35 USC 101 as being directed to an abstract idea without significantly more. Claim Rejections - 35 USC § 103 The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claims 2, 3, 5, 7-10, 12, 13, 15, and 17-20 are rejected under 35 U.S.C. 103 as being unpatentable over Jang (US 20140359523 A1), in view of Ramos et al. (US 11157696 B1), and in further view of Olstad et al. (US 20130132374 A1). With regard to claim 2, Jang teaches a method for responding to voice queries (Fig. 11; [0165]-[0169]; Fig. 12; [0170]-[0173]), the method comprising: receiving a voice query received at an audio interface ([0166]; Fig. 1, microphone 122; [0047]: receive a voice query through an audio input component such as a microphone, wherein the microphone corresponds to “an audio interface”); extracting, using control circuitry (Fig. 1; [0073]-[0076]: controller 180 corresponds to “control circuity”), one or more keywords based at least in part on the voice query ([0170]: identify a query term of the voice query, wherein a query term corresponds to “one or more keywords based at least in part on the voice query”); generating, using the control circuitry, a text query based at least in part on the text representation (Fig. 11; [0167]-[0168]; Fig. 12; [0170]: convert the voice query to a text query, which identifies query terms from the voice query, determines pronunciation information for each query term, and converts each query term into a typical text query term using a voice query term database that links a range of pronunciation of terms to a typical query term. As a result, a text query is generated comprising the query terms and their pronunciations); Jang does not explicitly teach selecting, using the control circuitry, a text representation corresponding to the one or more keywords based at least in part on one or more text representations accessed in a data structure of a user profile associated with the voice query; identifying a first entity of a plurality of entities based at least in part on a similarity of the text query to metadata stored for the first entity, wherein the metadata comprises one or more alternate text representations of an identifier associated with the first entity; and retrieving a content item associated with the first entity. Ramos teaches selecting, using the control circuitry, a text representation corresponding to the one or more keywords based at least in part on one or more text representations accessed in a data structure of a user profile associated with the voice query (due to the indefiniteness discussed above, this limitation is interpreted to mean using user profile information to help determine user audio queries. This is taught by claim 2); identifying a first entity of a plurality of entities based at least in part on a similarity of the text query to metadata stored for the first entity, wherein the metadata comprises a pronunciation tag comprising a phonetic spelling for the entity (Fig. 1, 138-144; Col. 3, lines 61-67; Col. 4, lines 1-56: perform entity resolution based on the tagged portion of text data and the portion of audio data corresponding to the tagged portion of text by comparing the portion of audio data against audio data representing entities known to the system, wherein performing entity resolution corresponds to “identifying a first entity”, the tagged portion of text data corresponds to “the text query”, and audio data representing entities known to the system corresponds to “a pronunciation tag” comprised in the “metadata stored for the entity”. Fig. 6; Col. 16, lines 42-67; Col. 17, lines 1-18: audio data comprises phonetic representation of text data, wherein the phonetic representation reads on "phonetic spelling"); and retrieving a content item associated with the first entity (Fig. 1, 146; Col. 4, lines 57-62; Col. 2, lines 10-18: use the resolved entity to perform downstream processes. For example, for the user input of "Alexa, play Adele music," a system may output music sung by Adele, wherein output indicates “retrieving”, and music sung by Adele corresponds to “a content item associated with the first entity” Adele). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified Jang to incorporate the teachings of Ramos to identify an entity based at least in part on the text query and metadata for the entity, wherein the metadata comprises one or more alternate text representations of an identifier associated with the entity and retrieve a content item associated with the entity. Doing so would improve text-based entity resolution by providing a language agnostic phonetic searching as part of entity resolution when text-based entity resolution may be unsuccessful or successful to a degree below a requisite threshold confidence as taught by Ramos (Col. 2, lines 48-67). Jang and Ramos do not teach identifying a first entity of a plurality of entities based at least in part on a similarity of the text query to metadata stored for the first entity, wherein the metadata comprises one or more alternate text representations of an identifier associated with the first entity; Olstad teaches identifying a first entity of a plurality of entities based at least in part on a similarity of the text query to metadata stored for the first entity, wherein the metadata comprises one or more alternate text representations of an identifier associated with the first entity (Claim 20; [0039]: identify videos in the set of videos based on metadata associated with the videos, wherein the metadata includes phonetic transcription extracted from the audio track, and the phoneme sequences included in a phonetic transcription of the audio track are matched with a phonetic representation of the query to find locations inside the audio track with the best phonetic similarity. “query terms” indicates “text query” and phonetic transcription metadata is “one or more alternate text representations” of a speech-to-text transcription associated with the identified videos, i.e., “an identifier”); It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified Jang and Ramos to incorporate the teachings of Olstad to identify an entity based at least in part on the text query and metadata for the entity, wherein the metadata comprises one or more alternate text representations of an identifier associated with the entity. Doing so would find locations inside the audio track with the best phonetic similarity to a user query, improve search precision, and perform less analysis including metadata generation as taught by Olstad ([0039]). With regard to claim 3, As discussed in claim 2, Jang and Ramos and Olstad teach all the limitations therein. Olstad further teaches the method of claim 2, wherein the one or more alternate text representations comprises a phonetic representation of the identifier associated with the first entity (Claim 20; [0039]: the phonetic transcription included in the metadata that is extracted from the audio track of the identified videos corresponds to “a phonetic representation”). With regard to claim 5, As discussed in claim 2, Jang and Ramos and Olstad teach all the limitations therein. Olstad further teaches the method of claim 2, wherein the one or more alternate text representations comprises a text string generated based at least in part on a previous speech-to-text conversion. ([0039]: phonetic transcription being an alternative to speech-to-text transcription of the audio track in the video indicates generating a text string by either speech-to-text conversion or phonetic transcription, which takes place before the metadata is used for search, i.e., “a previous”). With regard to claim 7, As discussed in claim 2, Jang and Ramos and Olstad teach all the limitations therein. Olstad further teaches the method of claim 2, wherein the identifier associated with the first entity identifies information related to the entity (Claim 20; [0039]: the phonetic transcription included in the metadata of the identified videos identifies the audio track in the videos, i.e., “identifies information related to” the videos, i.e., “the first entity”). With regard to claim 8, As discussed in claim 2, Jang and Ramos and Olstad teach all the limitations therein. Ramos further teaches the method of claim 2, wherein the identifying the first entity is based at least in part on user profile information. (Col. 20, lines 13-19; Col. 7, lines 61-66: the phonetic entity resolution component 802 may consider user preferences, wherein user preferences read on “user profile information”). With regard to claim 9, As discussed in claim 2, Jang and Ramos and Olstad teach all the limitations therein. Ramos further teaches the method of claim 2, wherein the identifying the first entity is based at least in part on popularity information associated with the entity. (Col. 20, lines 13-19: the phonetic entity resolution component 802 may consider popularity of known entities). With regard to claim 10, As discussed in claim 2, Jang and Ramos and Olstad teach all the limitations therein. Olstad further teaches the method of claim 2, further comprising: identifying a second entity based at least in part on the text query and metadata for the second entity (Claim 20; [0039]; [0003]: querying videos by matching a phonetic representation of the query to phonetic transcription of the audio track inside the videos indicates identifying plural videos that include a first video, a second video, and potentially more, wherein each video corresponds to an “entity”), and Ramos further teaches determining a first score for the first entity based at least in part on a comparison of the text query to the metadata associated with the first entity, and determining a second score for the second entity based at least in part on a comparison of the text query to metadata associated with the second entity, wherein the content item associated with the first entity is retrieved by selecting a maximum score of the first score and the second score (Fig. 8; Col. 19, lines 27-41: perform phonetic matching of the audio data (representing the entity to be resolved) to audio data stored in the entity storage (608/706) and associate with each known entity a confidence value representing the confidence that the known entity corresponds to the entity in the user input, wherein the confidence value corresponds to a “score” and the phonetic matching corresponds to a “comparison”. Col. 20, lines 20-41: the N-best list may include a maximum number of top scoring known entities, wherein selecting the maximum number of top scoring known entities indicates not only every entity is scored but also only the ones with "a maximum score" gets selected). With regard to claim 12, Jang teaches a system for responding to voice queries (Fig. 11; [0165]-[0169]; Fig. 12; [0170]-[0173]), the system comprising: a memory (Fig. 1: memory 160); and control circuitry (Fig. 1; [0073]-[0076]: controller 180 corresponds to “control circuity”) configured to: receive a voice query received at an audio interface ([0166]; Fig. 1, microphone 122; [0047]: receive a voice query through an audio input component such as a microphone, wherein the microphone corresponds to “an audio interface”); store in the memory the voice query (since the memory is the place for data a computer needs to access quickly for processing, receiving a voice query as discussed above in [0166]; Fig. 1, microphone 122; [0047] and processing the voice query to extract keywords as discussed below in [0170] both inherently teach storing the received voice query in the memory); extract one or more keywords based at least in part on the voice query ([0170]: identify a query term of the voice query, wherein a query term corresponds to “one or more keywords based at least in part on the voice query”); generate a text query based at least in part on the text representation (Fig. 11; [0167]-[0168]; Fig. 12; [0170]: convert the voice query to a text query, which identifies query terms from the voice query, determines pronunciation information for each query term, and converts each query term into a typical text query term using a voice query term database that links a range of pronunciation of terms to a typical query term. As a result, a text query is generated comprising the query terms and their pronunciations); Jang does not explicitly teach identify a first entity of a plurality of entities based at least in part on a similarity of the text query to metadata stored for the entity, wherein the metadata comprises one or more alternate text representations of an identifier associated with the first entity; and retrieve a content item associated with the first entity. Ramos teaches identify a first entity of a plurality of entities based at least in part on a similarity of the text query to metadata stored for the entity, wherein the metadata comprises a pronunciation tag comprising a phonetic spelling for the entity (Fig. 1, 138-144; Col. 3, lines 61-67; Col. 4, lines 1-56: perform entity resolution based on the tagged portion of text data and the portion of audio data corresponding to the tagged portion of text by comparing the portion of audio data against audio data representing entities known to the system, wherein performing entity resolution corresponds to “identifying a first entity”, the tagged portion of text data corresponds to “the text query”, and audio data representing entities known to the system corresponds to “a pronunciation tag” comprised in the “metadata stored for the entity”. Fig. 6; Col. 16, lines 42-67; Col. 17, lines 1-18: audio data comprises phonetic representation of text data, wherein the phonetic representation reads on "phonetic spelling"); and retrieve a content item associated with the first entity (Fig. 1, 146; Col. 4, lines 57-62; Col. 2, lines 10-18: use the resolved entity to perform downstream processes. For example, for the user input of "Alexa, play Adele music," a system may output music sung by Adele, wherein output indicates “retrieving”, and music sung by Adele corresponds to “a content item associated with the first entity” Adele). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified Jang to incorporate the teachings of Ramos to identify an entity based at least in part on the text query and metadata for the entity, wherein the metadata comprises one or more alternate text representations of an identifier associated with the entity and retrieve a content item associated with the entity. Doing so would improve text-based entity resolution by providing a language agnostic phonetic searching as part of entity resolution when text-based entity resolution may be unsuccessful or successful to a degree below a requisite threshold confidence as taught by Ramos (Col. 2, lines 48-67). Jang and Ramos do not teach identify a first entity of a plurality of entities based at least in part on a similarity of the text query to metadata stored for the entity, wherein the metadata comprises one or more alternate text representations of an identifier associated with the first entity; Olstad teaches identify a first entity of a plurality of entities based at least in part on a similarity of the text query to metadata stored for the entity, wherein the metadata comprises one or more alternate text representations of an identifier associated with the first entity (Claim 20; [0039]: identify videos in the set of videos based on metadata associated with the videos, wherein the metadata includes phonetic transcription extracted from the audio track, and the phoneme sequences included in a phonetic transcription of the audio track are matched with a phonetic representation of the query to find locations inside the audio track with the best phonetic similarity. “query terms” indicates “text query” and phonetic transcription metadata is “one or more alternate text representations” of a speech-to-text transcription associated with the identified videos, i.e., “an identifier”); It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified Jang and Ramos to incorporate the teachings of Olstad to identify an entity based at least in part on the text query and metadata for the entity, wherein the metadata comprises one or more alternate text representations of an identifier associated with the entity. Doing so would find locations inside the audio track with the best phonetic similarity to a user query, improve search precision, and perform less analysis including metadata generation as taught by Olstad ([0039]). With regard to claim 13, As discussed in claim 12, Jang and Ramos and Olstad teach all the limitations therein. Olstad further teaches the system of claim 12, wherein the one or more alternate text representations comprises a phonetic representation of the identifier associated with the first entity (Claim 20; [0039]: the phonetic transcription included in the metadata that is extracted from the audio track of the identified videos corresponds to “a phonetic representation”). With regard to claim 15, As discussed in claim 12, Jang and Ramos and Olstad teach all the limitations therein. Olstad further teaches the system of claim 12, wherein the one or more alternate text representations comprises a text string generated based at least in part on a previous speech-to-text conversion. ([0039]: phonetic transcription being an alternative to speech-to-text transcription of the audio track in the video indicates generating a text string by either speech-to-text conversion or phonetic transcription, which takes place before the metadata is used for search, i.e., “a previous”). With regard to claim 17, As discussed in claim 12, Jang and Ramos and Olstad teach all the limitations therein. Olstad further teaches the system of claim 12, wherein the identifier associated with the entity identifies information related to the first entity (Claim 20; [0039]: the phonetic transcription included in the metadata of the identified videos identifies the audio track in the videos, i.e., “identifies information related to” the videos, i.e., “the first entity”). With regard to claim 18, As discussed in claim 12, Jang and Ramos and Olstad teach all the limitations therein. Ramos further teaches the system of claim 12, wherein the identifying the first entity is based at least in part on user profile information. (Col. 20, lines 13-19; Col. 7, lines 61-66: the phonetic entity resolution component 802 may consider user preferences, wherein user preferences read on “user profile information”). With regard to claim 19, As discussed in claim 12, Jang and Ramos and Olstad teach all the limitations therein. Ramos further teaches the system of claim 12, wherein the identifying the entity is based at least in part on popularity information associated with the first entity. (Col. 20, lines 13-19: the phonetic entity resolution component 802 may consider popularity of known entities). With regard to claim 20, As discussed in claim 12, Jang and Ramos and Olstad teach all the limitations therein. Olstad further teaches the system of claim 12, the system is configured to: identify a second entity based at least in part on the text query and metadata for the second entity (Claim 20; [0039]; [0003]: querying videos by matching a phonetic representation of the query to phonetic transcription of the audio track inside the videos indicates identifying plural videos that include a first video, a second video, and potentially more, wherein each video corresponds to an “entity”), and Ramos further teaches determine a first score for the first entity based at least in part on a comparison of the text query to the metadata associated with the first entity, and determine a second score for the second entity based at least in part on a comparison of the text query to metadata associated with the second entity, wherein the content item associated with the first entity is retrieved by selecting a maximum score of the first score and the second score (Fig. 8; Col. 19, lines 27-41: perform phonetic matching of the audio data (representing the entity to be resolved) to audio data stored in the entity storage (608/706) and associate with each known entity a confidence value representing the confidence that the known entity corresponds to the entity in the user input, wherein the confidence value corresponds to a “score” and the phonetic matching corresponds to a “comparison”. Col. 20, lines 20-41: the N-best list may include a maximum number of top scoring known entities, wherein selecting the maximum number of top scoring known entities indicates not only every entity is scored but also only the ones with "a maximum score" gets selected). Claims 4 and 14 are rejected under 35 U.S.C. 103 as being unpatentable over Jang (US 20140359523 A1), in view of Ramos et al. (US 11157696 B1), and in further view of Olstad et al. (US 20130132374 A1) and YAO et al. (US 20110307432 A1). With regard to claim 4, As discussed in claim 2, Jang and Ramos and Olstad teach all the limitations therein. Jang and Ramos and Olstad do not teach the method of claim 2, wherein the one or more alternate text representations comprises an alternate spelling of the identifier associated with the first entity. YAO teaches the method of claim 2, wherein the one or more alternate text representations comprises an alternate spelling of the identifier associated with the first entity ([0055]: indexed metadata for web pages includes entity name equivalents data, variations of an entity's name, and entity name misspellings, all of which correspond to "an alternate spelling" associated with the web pages). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified Jang and Ramos and Olstad to incorporate the teachings of YAO to make the one or more alternate text representations comprise an alternate spelling of the identifier associated with the entity. Doing so would provide improved search result relevance for name search queries as taught by YAO ([0005]). With regard to claim 14, As discussed in claim 12, Jang and Ramos and Olstad teach all the limitations therein. Jang and Ramos and Olstad do not teach the system of claim 12, wherein the one or more alternate text representations comprises an alternate spelling of the identifier associated with the first entity. YAO teaches the system of claim 12, wherein the one or more alternate text representations comprises an alternate spelling of the identifier associated with the first entity ([0055]: indexed metadata for web pages includes entity name equivalents data, variations of an entity's name, and entity name misspellings, all of which correspond to "an alternate spelling" associated with the web pages). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified Jang and Ramos and Olstad to incorporate the teachings of YAO to make the one or more alternate text representations comprise an alternate spelling of the identifier associated with the entity. Doing so would provide improved search result relevance for name search queries as taught by YAO ([0005]). Claims 6 and 16 are rejected under 35 U.S.C. 103 as being unpatentable over Jang (US 20140359523 A1), in view of Ramos et al. (US 11157696 B1), and in further view of Olstad et al. (US 20130132374 A1) and Pore et al. (US 20190295527 A1). With regard to claim 6, As discussed in claim 2, Jang and Ramos and Olstad teach all the limitations therein. Jang and Ramos and Olstad do not teach the method of claim 2, wherein the metadata is generated by a text-to-speech module using a pronunciation setting for generating speech content; and converting the generated speech content into text using a speech-to-text module. Pore teaches the method of claim 2, wherein the metadata is generated by a text-to-speech module using a pronunciation setting for generating speech content; and converting the generated speech content into text using a speech-to-text module (Abstract; [0035]-[0036]: convert a text message containing at least one phonemic spelling of a word into speech by running a text-to-speech application programming interface (API) with the text message as input. The converted speech may be input to a speech-to-text API and the speech-to-text API executed to convert the speech to text. Users of English language in different geographic locations may have different accents or pronunciations, and therefore, may pronounce or voice an English word based on phonemes particular to the geographic location. A speech-to-text API that can recognize the particular location's accents or pronunciation of words may provide for a more accurate conversion of speech into text). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified Jang and Ramos and Olstad to incorporate the teachings of Pore to generate the metadata by a text-to-speech module using a pronunciation setting for generating speech content, and convert the generated speech content into text using a speech-to-text module. Doing so would recognize a particular location's accents of a user and provide more accurate conversion of the text into English speech as taught by Pore ([0035]). With regard to claim 16, As discussed in claim 12, Jang and Ramos and Olstad teach all the limitations therein. Jang and Ramos and Olstad do not teach the system of claim 12, wherein the metadata is generated by a text-to-speech module using a pronunciation setting for generating speech content; and the system is configured to: converting the generated speech content into text using a speech-to-text module. Pore teaches the system of claim 12, wherein the metadata is generated by a text-to-speech module using a pronunciation setting for generating speech content; and the system is configured to: converting the generated speech content into text using a speech-to-text module (Abstract; [0035]-[0036]: convert a text message containing at least one phonemic spelling of a word into speech by running a text-to-speech application programming interface (API) with the text message as input. The converted speech may be input to a speech-to-text API and the speech-to-text API executed to convert the speech to text. Users of English language in different geographic locations may have different accents or pronunciations, and therefore, may pronounce or voice an English word based on phonemes particular to the geographic location. A speech-to-text API that can recognize the particular location's accents or pronunciation of words may provide for a more accurate conversion of speech into text). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified Jang and Ramos and Olstad to incorporate the teachings of Pore to generate the metadata by a text-to-speech module using a pronunciation setting for generating speech content, and convert the generated speech content into text using a speech-to-text module. Doing so would recognize a particular location's accents of a user and provide more accurate conversion of the text into English speech as taught by Pore ([0035]). Claims 11 and 21 are rejected under 35 U.S.C. 103 as being unpatentable over Jang (US 20140359523 A1), in view of Ramos et al. (US 11157696 B1), and in further view of Olstad et al. (US 20130132374 A1) and Davallou (US 20060074892 A1). With regard to claim 11, As discussed in claim 2, Jang and Ramos and Olstad teach all the limitations therein. Jang and Ramos and Olstad do not teach the method of claim 2, wherein the text query is a first text query, and further comprising: generating a plurality of text queries, wherein the plurality of text queries comprises the first text query, and wherein each text query of the plurality of text queries is generated based at least in part on a respective pronunciation setting of a speech-to-text module. Davallou teaches the method of claim 2, wherein the text query is a first text query, and further comprising: generating a plurality of text queries, wherein the plurality of text queries comprises the first text query, and wherein each text query of the plurality of text queries is generated based at least in part on a respective pronunciation setting of a speech-to-text module ([0064]: generate the search string through a speech-to-text software application from the dictation of a user interacting with a database, such as through a telephone system or microphone. [0079]; [0069]: search for similar sounding words within the phonetic database for each word and generate a number of different combinations based on the approximate pronunciation of the text entered and the phonetically equivalent formulas, the result of which would be “generating a plurality of text queries” based on a pronunciation setting of a speech-to-text module). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified Jang and Ramos and Olstad to incorporate the teachings of Davallou to generate a plurality of text queries, wherein the plurality of text queries comprises the first text query, and wherein each text query of the plurality of text queries is generated based at least in part on a respective pronunciation setting of a speech-to-text module. Doing so would eventually find the correct query even though the original entry involved a different spelling, without the user having to respell or retype the entry as taught by Davallou ([0069]). With regard to claim 21, As discussed in claim 12, Jang and Ramos and Olstad teach all the limitations therein. Jang and Ramos and Olstad do not teach the system of claim 12, wherein the text query is a first text query, and further comprising: generating a plurality of text queries, wherein the plurality of text queries comprises the first text query, and wherein each text query of the plurality of text queries is generated based at least in part on a respective pronunciation setting of a speech-to-text module. Davallou teaches the system of claim 12, wherein the text query is a first text query, and further comprising: generating a plurality of text queries, wherein the plurality of text queries comprises the first text query, and wherein each text query of the plurality of text queries is generated based at least in part on a respective pronunciation setting of a speech-to-text module ([0064]: generate the search string through a speech-to-text software application from the dictation of a user interacting with a database, such as through a telephone system or microphone. [0079]; [0069]: search for similar sounding words within the phonetic database for each word and generate a number of different combinations based on the approximate pronunciation of the text entered and the phonetically equivalent formulas, the result of which would be “generating a plurality of text queries” based on a pronunciation setting of a speech-to-text module). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified Jang and Ramos and Olstad to incorporate the teachings of Davallou to generate a plurality of text queries, wherein the plurality of text queries comprises the first text query, and wherein each text query of the plurality of text queries is generated based at least in part on a respective pronunciation setting of a speech-to-text module. Doing so would eventually find the correct query even though the original entry involved a different spelling, without the user having to respell or retype the entry as taught by Davallou ([0069]). Conclusion Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a). A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any extension fee pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the date of this final action. Any inquiry concerning this communication or earlier communications from the examiner should be directed to XIAOQIN HU whose telephone number is (571)272-1792. The examiner can normally be reached on Monday-Friday 7:00am-3:30pm. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Charles Rones can be reached on (571) 272-4085. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of an application may be obtained from the Patent Application Information Retrieval (PAIR) system. Status information for published applications may be obtained from either Private PAIR or Public PAIR. Status information for unpublished applications is available through Private PAIR only. For more information about the PAIR system, see http://pair-direct.uspto.gov. Should you have questions on access to the Private PAIR system, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative or access to the automated information system, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /XIAOQIN HU/Examiner, Art Unit 2168 /CHARLES RONES/Supervisory Patent Examiner, Art Unit 2168
Read full office action

Prosecution Timeline

May 19, 2025
Application Filed
Mar 11, 2026
Non-Final Rejection mailed — §101, §103, §112
Aug 11, 2026
Response Filed
Sep 23, 2026
Final Rejection mailed — §101, §103, §112 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12724834
SYSTEM AND METHOD FOR AUTOMATED FILE REPORTING
1y 6m to grant Granted Sep 01, 2026
Patent 12681923
VISUALLY MAPPING NODES AND CONNECTIONS IN ONE OR MORE ENTERPRISE-LEVEL SYSTEMS
1y 7m to grant Granted Jul 14, 2026
Patent 12670173
AUTOMATED EXTRACT, TRANSFORM, AND LOAD PROCESS
1y 7m to grant Granted Jun 30, 2026
Patent 12608383
BULK MATCHING DATA RECORD ENTITIES
2y 6m to grant Granted Apr 21, 2026
Patent 12585863
COMPRESSION SCHEME FOR STABLE UNIVERSAL UNIQUE IDENTITIES
1y 3m to grant Granted Mar 24, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

3-4
Expected OA Rounds
62%
Grant Probability
99%
With Interview (+56.2%)
2y 10m (~1y 6m remaining)
Median Time to Grant
Moderate
PTA Risk
Based on 195 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month