DETAILED ACTION
Continued Examination Under 37 CFR 1.114
A request for continued examination under 37 CFR 1.114, including the fee set forth in 37 CFR 1.17(e), was filed in this application after final rejection. Since this application is eligible for continued examination under 37 CFR 1.114, and the fee set forth in 37 CFR 1.17(e) has been timely paid, the finality of the previous Office action has been withdrawn pursuant to 37 CFR 1.114. Applicant's submission filed on 06/04/2026 has been entered.
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Claims 1 - 22 are pending and claims 1, 8 and 15 are independent claims.
Response to Arguments
Applicant’s arguments, see pages 8-11, filed on 06/04/2026, with respect to the 35 U.S.C. 101 rejections have been fully considered and they are persuasive.
The Applicant has made a request of confirmation on their interpretation of claims 9 and 21 rejection, esp. on the Final Office Action, regarding references Garg and Gupta. The Applicant states that in the rejections of claims 9 and 21, the Office action cites to both Gupta and Garg and the Applicant has interpreted the references to "Garg" to actually refer to Gupta (Arguments, page 8).
The Examiner would like to confirm and clarify any confusion of the Garg reference with the Gupta reference. The Examiner was responding to the Applicant’s argument about Elisha and Garg based on the first Nonfinal rejection. The Examiner used only Elisha and Garg to reject all the independent claims and most of the dependent claims on that first Nonfinal. However, the Final rejection was based on the Applicant’s amended claims which were then rejected with a combination of Elisha, Gupta and Jia, as the Applicant clearly stated. This is, thus, to confirm that the Applicant’s interpretation of the reference "Garg" to actually refer to Gupta in the rejections of claims 9 and 21 was correct.
Applicant’s arguments, see pages 11-17, filed on 06/04/2026, with respect to the 35 U.S.C. 103 rejections of claims 1-22 have been fully considered but are moot because the new ground of rejection does not rely on any reference applied in the prior rejection of record for any teaching or matter specifically challenged in the argument.
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claims 1, 3, 6-8, 10, 13-15, 17 and 20-21 are rejected under 35 U.S.C. 103 as being unpatentable over Kumaran, A. "MIRA: Multilingual information processing on relational architecture," In International Conference on Extending Database Technology, pp. 12-23, Berlin, Heidelberg: Springer Berlin Heidelberg, 2004 (Kumaran) in view of Gupta et al., "Query expansion for mixed-script information retrieval, SIGIR '14: Proceedings of the 37th international ACM SIGIR conference on Research & development in information retrieval, 2014 (Gupta), and further in view of Jia et al. Pat App No. US 20220310059 A1 (Jia).
Regarding Claim 1, Kumaran discloses a method of performing a cross-lingual search of an inverted index database for an input token of a search query (Kumaran, page 12, 1st para, Our proposal, Multilingual Information processing on Relational Architecture (MIRA), attempts to enhance the relational database systems with multilingual features and to make the query performance nearly language neutral; Kumaran, page 16, 2nd para, In this section we briefly sketch our phonetic matching approach that extends earlier works in monolingual world to matching of multilingual names. We further enhance the performance of such matching by defining a phonemic index based on the classic Soundex algorithm [10] and by adopting q-gram techniques [7] that has been successfully used in approximate matching of monolingual names to multilingual world ), the method comprising:
generating a phonemic index of the input token in a first orthography of a first language (Kumaran, page 22, 2nd para, Several approximate indexing methodologies offer search capability on pre-generated phonemic strings corresponding to names; [i.e., “names” as “input token”] ) by converting the input token into a phonetic representation and transcribing the phonetic representation of the input token into the phonemic index, wherein the phonemic index includes phoneme embeddings corresponding to each phoneme of the input token (Kumaran, page 16, 3rd para – page 17, para 1, We propose a phonemic matching strategy (shown as dotted line inFigure4) in LexEQUAL operator, as follows: First, the multilingual text strings are transformed to their equivalent phonemic representations in International Phonetic Alphabet (IPA)4 [3], obtained using standard text-to-phoneme (TTP) converters. The resulting phoneme strings represent a normalized form of proper names across languages, thus providing a means of comparison. Further, when the text data is stored in multiple scripts, this may be the only means of comparing them);
executing an approximate matching analysis on content tokens of the inverted index database based on the phonemic index of the input token by retrieving phonemic indices of the content tokens from a first inverted index and a second inverted index and (Kumaran, page 12, 2nd para, our proposed architecture is amenable for easy implementation in any type of query processing and information retrieval systems; page 13, 4th para – 14th page, 2nd para, In this environment, suppose a user wants to search for the works of an author in all (or a specified set of) languages… We refer matching on multi-lexical text strings, based on their phonemic equivalence as Multi-lexical Phonemic Matching. Though restricted to proper names, such matching represent a significant part of the user query strings in text databases and search engines, as proper and generic names constitute a fifth of normal corpora [9] ),
returning one or more search results based on the approximate matching analysis, the one or more search results including results associated with the first orthography and results associated with the second orthography (Kumaran, 13th page, 5th para - 14th page, 1st para, A sample phonetic query and the corresponding answer set, when issued on Books.com, are given in Figure 2. The returned tuples have in Author column the multi lexical strings that are phonemically close to the query string in English, namely, Nehru. The specification of ALL for the list of languages would have brought all records containing author names that are phonetically equivalent to Nehru, irrespective of the languages).
Kumaran does not specifically disclose wherein the inverted index database include the first inverted index corresponding to phonemic variants of the content tokens and to the first orthography and the second inverted index corresponding to phonemic variants of the content tokens and to a second orthography of a second language.
However, Gupta, in the same field of endeavor, discloses wherein the inverted index database includes the first inverted index corresponding to phonemic variants of the content tokens and to the first orthography and the second inverted index corresponding to phonemic variants of the content tokens and to a second orthography of a second language (Gupta, page 682, col 2, 2nd para – page 683, 1st col, 3rd para, Now we describe the experimental set up for evaluating the effectiveness of the proposed method for retrieval in Mixed-Script space. 5.1 Dataset We used the FIRE 2013 shared task collection on Transliterated Search [26] for experiments and training. The dataset comprises of document collection, queryset (Q) and relevance judgments. The collection (D1) contains 62,888 documents containing song title and lyrics in Roman, Devanagari and mixed scripts. Statistics of the document collection is given in Table 2 (a). The Q contains 25 lyrics search queries for Bollywood songs in Roman script with mean query length of 4.5 words. Table 2 (b) lists a few examples of queries from Q.
PNG
media_image1.png
96
348
media_image1.png
Greyscale
The experimental setup is a standard adhoc retrieval setting. The document collection is first indexed to create an inverted index and the index lexicon is used as mining lexicon. Being this a lyrics retrieval set up, the sequential information among the terms is crucial for effectiveness evaluation, e.g. “love me baby” and “baby love me” are completely different songs. In order to capture the word-ordering we consider word 2-grams as a unit for indexing and retrieval. The non-trivial part of mixed-script IR is query-enrichment to handle the challenges described in Sec. 2. In order to enrich the query with equivalents, we find the equivalents of the query terms as described in Section 4.5 and the word 2-gram query is formulated as shown in [13]. We consider a variety of systems to be compared with the proposed method. The query formulation is similar for all the systems including the retrieval settings like inverted index, retrieval model and mining lexicon except the method of finding the equivalents… The problem of finding equivalents is formulated as searching across the views by learning hashing functions as presented in [20]… An inverted index of hash codes is prepared for terms in mining lexicon. The equivalents for the query term are found from this index according to the score given by the graph matching algorithm (according to the cosine similarity in the common geometric space) of [30]).
Therefore, it would have been obvious for one having ordinary skill in the art before the effective filing date of the claimed invention to incorporate the method of Gupta in the method of Kumaran because this would present an extensive empirical analysis of the proposed method along with the evaluation results in an ad-hoc retrieval setting of mixedscript IR where the proposed method achieves significantly better results (12% increase in MRR and 29% increase in MAP) compared to other state-of-the-art baselines(Gupta, Abstract).
Kumaran in view of Gupta do not disclose comparing phoneme embeddings corresponding to each phoneme of the input token to phoneme embeddings of each phoneme of each content token.
However, Jia, in the same field of endeavor, discloses by comparing phoneme embeddings corresponding to each phoneme of the input token to phoneme embeddings of each phoneme of each content token (Jia, para 0013, determining the respective grapheme token representing the respective word of sequence of words corresponding to the respective phoneme token may include determining that the respective grapheme token includes a corresponding word position embedding that matches the respective word position embedding of the respective phoneme token; [i.e., “grapheme token” as “input token” and “word position embedding of the respective phoneme token” as “content token”]).
Therefore, it would have been obvious for one having ordinary skill in the art before the effective filing date of the claimed invention to incorporate the method of Jia in the method of Kumaran in view of Gupta because this would enable the search engine return one or more search results that the device interprets to generate a response for the user (Jia, para 0036).
Regarding Claim 3, Kumaran in view of Gupta, and further in view of Jia discloses the method of claim 1, further comprising:
generating a phonemic index of a first content token corresponding to the first orthography (Kumaran, 16th page, 2nd para, defining a phonemic index based on the classic Soundex algorithm [10] and by adopting q-gram techniques [7] that has been successfully used in approximate matching of monolingual names to multilingual world ); and
Gupta further teaches:
adding the phonemic index of the first content token in the first orthography to an inverted index corresponding to the first orthography in the inverted index database (Gupta, page 682, col 2, 2nd para – page 683, 1st col, 3rd para, Now we describe the experimental set up for evaluating the effectiveness of the proposed method for retrieval in Mixed-Script space. 5.1 Dataset We used the FIRE 2013 shared task collection on Transliterated Search [26] for experiments and training. The dataset comprises of document collection, queryset (Q) and relevance judgments. The collection (D1) contains 62,888 documents containing song title and lyrics in Roman, Devanagari and mixed scripts. Statistics of the document collection is given in Table 2 (a). The Q contains 25 lyrics search queries for Bollywood songs in Roman script with mean query length of 4.5 words. Table 2 (b) lists a few examples of queries from Q.
PNG
media_image1.png
96
348
media_image1.png
Greyscale
The experimental setup is a standard adhoc retrieval setting. The document collection is first indexed to create an inverted index and the index lexicon is used as mining lexicon. Being this a lyrics retrieval set up, the sequential information among the terms is crucial for effectiveness evaluation, e.g. “love me baby” and “baby love me” are completely different songs. In order to capture the word-ordering we consider word 2-grams as a unit for indexing and retrieval. The non-trivial part of mixed-script IR is query-enrichment to handle the challenges described in Sec. 2. In order to enrich the query with equivalents, we find the equivalents of the query terms as described in Section 4.5 and the word 2-gram query is formulated as shown in [13]. We consider a variety of systems to be compared with the proposed method. The query formulation is similar for all the systems including the retrieval settings like inverted index, retrieval model and mining lexicon except the method of finding the equivalents… The problem of finding equivalents is formulated as searching across the views by learning hashing functions as presented in [20]… An inverted index of hash codes is prepared for terms in mining lexicon. The equivalents for the query term are found from this index according to the score given by the graph matching algorithm (according to the cosine similarity in the common geometric space) of [30] ).
Regarding Claim 6, Kumaran in view of Gupta and further in view of Jia disclose the method of claim 1, further comprising:
generating the phonemic index of a phonemic variant of a first content token corresponding to the second orthography (Kumaran, 16th page, 2nd para, defining a phonemic index based on the classic Soundex algorithm [10] and by adopting q-gram techniques [7] that has been successfully used in approximate matching of monolingual names to multilingual world ); and
Gupta further teaches:
adding the phonemic index of the phonemic variant of the first content token corresponding to the second orthography to an inverted index corresponding to the second orthography in the inverted index database (Gupta, page 682, col 2, 2nd para – page 683, 1st col, 3rd para, Now we describe the experimental set up for evaluating the effectiveness of the proposed method for retrieval in Mixed-Script space. 5.1 Dataset We used the FIRE 2013 shared task collection on Transliterated Search [26] for experiments and training. The dataset comprises of document collection, queryset (Q) and relevance judgments. The collection (D1) contains 62,888 documents containing song title and lyrics in Roman, Devanagari and mixed scripts. Statistics of the document collection is given in Table 2 (a). The Q contains 25 lyrics search queries for Bollywood songs in Roman script with mean query length of 4.5 words. Table 2 (b) lists a few examples of queries from Q.
PNG
media_image1.png
96
348
media_image1.png
Greyscale
The experimental setup is a standard adhoc retrieval setting. The document collection is first indexed to create an inverted index and the index lexicon is used as mining lexicon. Being this a lyrics retrieval set up, the sequential information among the terms is crucial for effectiveness evaluation, e.g. “love me baby” and “baby love me” are completely different songs. In order to capture the word-ordering we consider word 2-grams as a unit for indexing and retrieval. The non-trivial part of mixed-script IR is query-enrichment to handle the challenges described in Sec. 2. In order to enrich the query with equivalents, we find the equivalents of the query terms as described in Section 4.5 and the word 2-gram query is formulated as shown in [13]. We consider a variety of systems to be compared with the proposed method. The query formulation is similar for all the systems including the retrieval settings like inverted index, retrieval model and mining lexicon except the method of finding the equivalents… The problem of finding equivalents is formulated as searching across the views by learning hashing functions as presented in [20]… An inverted index of hash codes is prepared for terms in mining lexicon. The equivalents for the query term are found from this index according to the score given by the graph matching algorithm (according to the cosine similarity in the common geometric space) of [30]).
Regarding Claim 7, Kumaran in view of Gupta, and further in view of Jia disclose the method of claim 1, wherein generating the phonemic index of the input token comprises:
converting the input token into a phonetic representation (Kumaran, pages 16, 3rd para – page 17, 1st para, the multilingual text strings are transformed to their equivalent phonemic representations ); and
transcribing the phonetic representation of the input token into the phonemic index of the input token, wherein the phonemic index includes phoneme embeddings corresponding to each phoneme of the input token (Kumaran, page 16, 3rd para – page 17, para 1, We propose a phonemic matching strategy (shown as dotted line inFigure4) in LexEQUAL operator, as follows: First, the multilingual text strings are transformed to their equivalent phonemic representations in International Phonetic Alphabet (IPA)4 [3], obtained using standard text-to-phoneme (TTP) converters. The resulting phoneme strings represent a normalized form of proper names across languages, thus providing a means of comparison. Further, when the text data is stored in multiple scripts, this may be the only means of comparing them ).
Regarding Claim 8, Kumaran discloses a computing system for searching an inverted index database for an input token of a search query (Kumaran, page 12, 1st para, Our proposal, Multilingual Information processing on Relational Architecture (MIRA), attempts to enhance the relational database systems with multilingual features and to make the query performance nearly language neutral; Kumaran, page 16, 2nd para, In this section we briefly sketch our phonetic matching approach that extends earlier works in monolingual world to matching of multilingual names. We further enhance the performance of such matching by defining a phonemic index based on the classic Soundex algorithm [10] and by adopting q-gram techniques [7] that has been successfully used in approximate matching of monolingual names to multilingual world ), the computing system comprising:
a phonemic indexer executable by the one or more hardware processors and configured to generate a phonemic index of the input token of a first orthography of a first language (Kumaran, page 22, 2nd para, Several approximate indexing methodologies offer search capability on pre-generated phonemic strings corresponding to names; [i.e., “names” as “input token”]) by converting the input token into a phonetic representation and transcribing the phonetic representation of the input token into the phonemic index, wherein the phonemic index includes phoneme embeddings corresponding to each phoneme of the input token (Kumaran, page 16, 3rd para – page 17, para 1, We propose a phonemic matching strategy (shown as dotted line inFigure4) in LexEQUAL operator, as follows: First, the multilingual text strings are transformed to their equivalent phonemic representations in International Phonetic Alphabet (IPA)4 [3], obtained using standard text-to-phoneme (TTP) converters. The resulting phoneme strings represent a normalized form of proper names across languages, thus providing a means of comparison. Further, when the text data is stored in multiple scripts, this may be the only means of comparing them);
an approximate match analyzer executable by the one or more hardware processors and configured to execute an approximate matching analysis on content tokens of the inverted index database based on the phonemic index of the input token by retrieving phonemic indices of the content tokens from a first inverted index and a second inverted index and (Kumaran, page 12, 2nd para, our proposed architecture is amenable for easy implementation in any type of query processing and information retrieval systems; page 13, 4th para – 14th page, 2nd para, In this environment, suppose a user wants to search for the works of an author in all (or a specified set of) languages… We refer matching on multi-lexical text strings, based on their phonemic equivalence as Multi-lexical Phonemic Matching. Though restricted to proper names, such matching represent a significant part of the user query strings in text databases and search engines, as proper and generic names constitute a fifth of normal corpora [9] ),
a score conditioner executable by the one or more hardware processors and configured to return one or more search results based on the approximate matching analysis, the one or more search results including results associated with the first orthography and results associated with the second orthography (Kumaran, 13th page, 5th para - 14th page, 1st para, A sample phonetic query and the corresponding answer set, when issued on Books.com, are given in Figure 2. The returned tuples have in Author column the multi lexical strings that are phonemically close to the query string in English, namely, Nehru. The specification of ALL for the list of languages would have brought all records containing author names that are phonetically equivalent to Nehru, irrespective of the languages ).
Kumaran does not specifically disclose wherein the inverted index database includes the first inverted index corresponding to phonemic variants of the content tokens and to the first orthography and the second inverted index corresponding to phonemic variants of the content tokens and to a second orthography of a second language; and
However, Gupta, in the same field of endeavor, discloses wherein the inverted index database includes the first inverted index corresponding to phonemic variants of the content tokens and to the first orthography and the second inverted index corresponding to phonemic variants of the content tokens and to a second orthography of a second language (Gupta, page 682, col 2, 2nd para – page 683, 1st col, 3rd para, Now we describe the experimental set up for evaluating the effectiveness of the proposed method for retrieval in Mixed-Script space. 5.1 Dataset We used the FIRE 2013 shared task collection on Transliterated Search [26] for experiments and training. The dataset comprises of document collection, queryset (Q) and relevance judgments. The collection (D1) contains 62,888 documents containing song title and lyrics in Roman, Devanagari and mixed scripts. Statistics of the document collection is given in Table 2 (a). The Q contains 25 lyrics search queries for Bollywood songs in Roman script with mean query length of 4.5 words. Table 2 (b) lists a few examples of queries from Q.
PNG
media_image1.png
96
348
media_image1.png
Greyscale
The experimental setup is a standard adhoc retrieval setting.The document collection is first indexed to create an inverted index and the index lexicon is used as mining lexicon. Being this a lyrics retrieval set up, the sequential information among the terms is crucial for effectiveness evaluation, e.g. “love me baby” and “baby love me” are completely different songs. In order to capture the word-ordering we consider word 2-grams as a unit for indexing and retrieval. The non-trivial part of mixed-script IR is query-enrichment to handle the challenges described in Sec. 2. In order to enrich the query with equivalents, we find the equivalents of the query terms as described in Section 4.5 and the word 2-gram query is formulated as shown in [13]. We consider a variety of systems to be compared with the proposed method. The query formulation is similar for all the systems including the retrieval settings like inverted index, retrieval model and mining lexicon except the method of finding the equivalents… The problem of finding equWe ivalents is formulated as searching across the views by learning hashing functions as presented in [20]… An inverted index of hashcodes is prepared for terms in mining lexicon. The equivalents for the query term are found from this index according to the score given by the graph matching algorithm (according to the cosine similarity in the common geometric space) of [30]).
Therefore, it would have been obvious for one having ordinary skill in the art before the effective filing date of the claimed invention to incorporate the method of Gupta in the method of Elisha because this would present an extensive empirical analysis of the proposed method along with the evaluation results in an ad-hoc retrieval setting of mixedscript IR where the proposed method achieves significantly better results (12% increase in MRR and 29% increase in MAP) compared to other state-of-the-art baselines(Gupta, Abstract).
Kumaran in view of Gupta do not disclose one or more hardware processors, and comparing phoneme embeddings corresponding to each phoneme of the input token to phoneme embeddings of each phoneme of each content token.
However, Jia, in the same field of endeavor, discloses:
one or more hardware processors (Jia, para 0053, The computing device 500 includes a processor 510 (e.g., data processing hardware), memory 520 (e.g., memory hardware), a storage device 530, a high-speed interface/controller 540 connecting to the memory 520 and high-speed expansion ports 550 );
comparing phoneme embeddings corresponding to each phoneme of the input token to phoneme embeddings of each phoneme of each content token Jia, para 0013, determining the respective grapheme token representing the respective word of sequence of words corresponding to the respective phoneme token may include determining that the respective grapheme token includes a corresponding word position embedding that matches the respective word position embedding of the respective phoneme token; [i.e., “grapheme token” as “input token” and “word position embedding of the respective phoneme token” as “content token”]).
Therefore, it would have been obvious for one having ordinary skill in the art before the effective filing date of the claimed invention to incorporate the method of Jia in the method of Kumaran in view of Gupta because this would enable the search engine return one or more search results that the device interprets to generate a response for the user (Jia, para 0036).
Regarding Claim 10, Kumaran in view of Gupta, and further in view of Jia disclose the computing system of claim 8, wherein the phonemic indexer is further configured to generate a phonemic index a first content token corresponding to the first orthography (Kumaran, 16th page, 2nd para, defining a phonemic index based on the classic Soundex algorithm [10] and by adopting q-gram techniques [7] that has been successfully used in approximate matching of monolingual names to multilingual world) and
Gupta further teaches:
to add the phonemic index the first content token in the first orthography to an inverted index corresponding to the first orthography in the inverted index database (Gupta, page 682, col 2, 2nd para – page 683, 1st col, 3rd para, Now we describe the experimental set up for evaluating the effectiveness of the proposed method for retrieval in Mixed-Script space. 5.1 Dataset We used the FIRE 2013 shared task collection on Transliterated Search [26] for experiments and training. The dataset comprises of document collection, queryset (Q) and relevance judgments. The collection (D1) contains 62,888 documents containing song title and lyrics in Roman, Devanagari and mixed scripts. Statistics of the document collection is given in Table 2 (a). The Q contains 25 lyrics search queries for Bollywood songs in Roman script with mean query length of 4.5 words. Table 2 (b) lists a few examples of queries from Q.
PNG
media_image1.png
96
348
media_image1.png
Greyscale
The experimental setup is a standard adhoc retrieval setting. The document collection is first indexed to create an inverted index and the index lexicon is used as mining lexicon. Being this a lyrics retrieval set up, the sequential information among the terms is crucial for effectiveness evaluation, e.g. “love me baby” and “baby love me” are completely different songs. In order to capture the word-ordering we consider word 2-grams as a unit for indexing and retrieval. The non-trivial part of mixed-script IR is query-enrichment to handle the challenges described in Sec. 2. In order to enrich the query with equivalents, we find the equivalents of the query terms as described in Section 4.5 and the word 2-gram query is formulated as shown in [13]. We consider a variety of systems to be compared with the proposed method. The query formulation is similar for all the systems including the retrieval settings like inverted index, retrieval model and mining lexicon except the method of finding the equivalents… The problem of finding equivalents is formulated as searching across the views by learning hashing functions as presented in [20]… An inverted index of hash codes is prepared for terms in mining lexicon. The equivalents for the query term are found from this index according to the score given by the graph matching algorithm (according to the cosine similarity in the common geometric space) of [30] ).
Regarding Claim 13, Kumaran in view of Gupta, and further in view of Jia disclose the computing system of claim 8, wherein the phonemic indexer is further configured to generate the phonemic index [[for]] a phonemic variant of a first content token corresponding to the second orthography (Kumaran, 16th page, 2nd para, defining a phonemic index based on the classic Soundex algorithm [10] and by adopting q-gram techniques [7] that has been successfully used in approximate matching of monolingual names to multilingual world) and
Gupta further teaches:
to add the phonemic index of the phonemic variant the first content token corresponding to the second orthography to an inverted index corresponding to the second orthography in the inverted index database (Gupta, page 682, col 2, 2nd para – page 683, 1st col, 3rd para, Now we describe the experimental set up for evaluating the effectiveness of the proposed method for retrieval in Mixed-Script space. 5.1 Dataset We used the FIRE 2013 shared task collection on Transliterated Search [26] for experiments and training. The dataset comprises of document collection, queryset (Q) and relevance judgments. The collection (D1) contains 62,888 documents containing song title and lyrics in Roman, Devanagari and mixed scripts. Statistics of the document collection is given in Table 2 (a). The Q contains 25 lyrics search queries for Bollywood songs in Roman script with mean query length of 4.5 words. Table 2 (b) lists a few examples of queries from Q.
PNG
media_image1.png
96
348
media_image1.png
Greyscale
The experimental setup is a standard adhoc retrieval setting. The document collection is first indexed to create an inverted index and the index lexicon is used as mining lexicon. Being this a lyrics retrieval set up, the sequential information among the terms is crucial for effectiveness evaluation, e.g. “love me baby” and “baby love me” are completely different songs. In order to capture the word-ordering we consider word 2-grams as a unit for indexing and retrieval. The non-trivial part of mixed-script IR is query-enrichment to handle the challenges described in Sec. 2. In order to enrich the query with equivalents, we find the equivalents of the query terms as described in Section 4.5 and the word 2-gram query is formulated as shown in [13]. We consider a variety of systems to be compared with the proposed method. The query formulation is similar for all the systems including the retrieval settings like inverted index, retrieval model and mining lexicon except the method of finding the equivalents… The problem of finding equivalents is formulated as searching across the views by learning hashing functions as presented in [20]… An inverted index of hash codes is prepared for terms in mining lexicon. The equivalents for the query term are found from this index according to the score given by the graph matching algorithm (according to the cosine similarity in the common geometric space) of [30]).
Regarding Claim 14, Kumaran in view of Gupta, and further in view of Jia disclose the computing system of claim 8, further comprising:
a phonetic converter executable by the one or more hardware processors and configured to convert the input token into a phonetic representation (Kumaran, pages 16, 3rd para – page 17, 1st para, the multilingual text strings are transformed to their equivalent phonemic representations ), wherein the phonemic indexer is further configured to transcribe the phonetic representation of the input token into the phonemic index page 16, 3rd para – page 17, para 1, We propose a phonemic matching strategy (shown as dotted line inFigure4) in LexEQUAL operator, as follows: First, the multilingual text strings are transformed to their equivalent phonemic representations in International Phonetic Alphabet (IPA)4 [3], obtained using standard text-to-phoneme (TTP) converters. The resulting phoneme strings represent a normalized form of proper names across languages, thus providing a means of comparison. Further, when the text data is stored in multiple scripts, this may be the only means of comparing them).
Regarding Claim 15, Kumaran discloses st para, Our proposal, Multilingual Information processing on Relational Architecture (MIRA), attempts to enhance the relational database systems with multilingual features and to make the query performance nearly language neutral; Kumaran, page 16, 2nd para, In this section we briefly sketch our phonetic matching approach that extends earlier works in monolingual world to matching of multilingual names. We further enhance the performance of such matching by defining a phonemic index based on the classic Soundex algorithm [10] and by adopting q-gram techniques [7] that has been successfully used in approximate matching of monolingual names to multilingual world ), the process comprising:
generating a phonemic index the input token of a first orthography of a first language (Kumaran, page 22, 2nd para, Several approximate indexing methodologies offer search capability on pre-generated phonemic strings corresponding to names; [i.e., “names” as “input token”] ) by converting the input token into a phonetic representation and transcribing the phonetic representation of the input token into the phonemic index, wherein the phonemic index includes phoneme embeddings corresponding to each phoneme of the input token (Kumaran, page 16, 3rd para – page 17, para 1, We propose a phonemic matching strategy (shown as dotted line inFigure4) in LexEQUAL operator, as follows: First, the multilingual text strings are transformed to their equivalent phonemic representations in International Phonetic Alphabet (IPA)4 [3], obtained using standard text-to-phoneme (TTP) converters. The resulting phoneme strings represent a normalized form of proper names across languages, thus providing a means of comparison. Further, when the text data is stored in multiple scripts, this may be the only means of comparing them);
executing an approximate matching analysis on content tokens of the inverted index database based on the phonemic index the input token by retrieving phonemic indices of the content tokens from a first inverted index and a second inverted index and (Kumaran, page 12, 2nd para, our proposed architecture is amenable for easy implementation in any type of query processing and information retrieval systems; page 13, 4th para – 14th page, 2nd para, In this environment, suppose a user wants to search for the works of an author in all (or a specified set of) languages… We refer matching on multi-lexical text strings, based on their phonemic equivalence as Multi-lexical Phonemic Matching. Though restricted to proper names, such matching represent a significant part of the user query strings in text databases and search engines, as proper and generic names constitute a fifth of normal corpora [9].),
returning one or more search results based on the approximate matching analysis, the one or more search results including results associated with the first orthography and results associated with the second orthography (Kumaran, 13th page, 5th para - 14th page, 1st para, A sample phonetic query and the corresponding answer set, when issued on Books.com, are given in Figure 2. The returned tuples have in Author column the multi lexical strings that are phonemically close to the query string in English, namely, Nehru. The specification of ALL for the list of languages would have brought all records containing author names that are phonetically equivalent to Nehru, irrespective of the languages )
Kumaran does not specifically disclose wherein the inverted index database includes the first inverted index corresponding to phonemic variants of the content tokens and to the first orthography and the second inverted index corresponding to phonemic variants of the content tokens and to a second orthography of a second language.
However, Gupta, in the same field of endeavor, discloses wherein the inverted index database includes the first inverted index corresponding to phonemic variants of the content tokens and to the first orthography and the second inverted index corresponding to phonemic variants of the content tokens and to a second orthography of a second language (Gupta, page 682, col 2, 2nd para – page 683, 1st col, 3rd para, Now we describe the experimental set up for evaluating the effectiveness of the proposed method for retrieval in Mixed-Script space. 5.1 Dataset We used the FIRE 2013 shared task collection on Transliterated Search [26] for experiments and training. The dataset comprises of document collection, queryset (Q) and relevance judgments. The collection (D1) contains 62,888 documents containing song title and lyrics in Roman, Devanagari and mixed scripts. Statistics of the document collection is given in Table 2 (a). The Q contains 25 lyrics search queries for Bollywood songs in Roman script with mean query length of 4.5 words. Table 2 (b) lists a few examples of queries from Q.
PNG
media_image1.png
96
348
media_image1.png
Greyscale
The experimental setup is a standard adhoc retrieval setting. The document collection is first indexed to create an inverted index and the index lexicon is used as mining lexicon. Being this a lyrics retrieval set up, the sequential information among the terms is crucial for effectiveness evaluation, e.g. “love me baby” and “baby love me” are completely different songs. In order to capture the word-ordering we consider word 2-grams as a unit for indexing and retrieval. The non-trivial part of mixed-script IR is query-enrichment to handle the challenges described in Sec. 2. In order to enrich the query with equivalents, we find the equivalents of the query terms as described in Section 4.5 and the word 2-gram query is formulated as shown in [13]. We consider a variety of systems to be compared with the proposed method. The query formulation is similar for all the systems including the retrieval settings like inverted index, retrieval model and mining lexicon except the method of finding the equivalents… The problem of finding equivalents is formulated as searching across the views by learning hashing functions as presented in [20]… An inverted index of hashcodes is prepared for terms in mining lexicon. The equivalents for the query term are found from this index according to the score given by the graph matching algorithm (according to the cosine similarity in the common geometric space) of [30])).
Therefore, it would have been obvious for one having ordinary skill in the art before the effective filing date of the claimed invention to incorporate the method of Gupta in the method of Kumaran because this would present an extensive empirical analysis of the proposed method along with the evaluation results in an ad-hoc retrieval setting of mixedscript IR where the proposed method achieves significantly better results (12% increase in MRR and 29% increase in MAP) compared to other state-of-the-art baselines(Gupta, Abstract).
Kumaran in view of Gupta do not disclose one or more tangible processor-readable storage media embodied with instructions for executing on one or more processors and circuits of a computing device, and comparing phoneme embeddings corresponding to each phoneme of the input token to phoneme embeddings of each phoneme of each content token.
However, Jia, in the same field of endeavor, discloses:
one or more tangible processor-readable storage media embodied with instructions for executing on one or more processors and circuits of a computing device (Jia, para 0053, The computing device 500 includes a processor 510 (e.g., data processing hardware), memory 520 (e.g., memory hardware), a storage device 530, a high-speed interface/controller 540 connecting to the memory 520 and high-speed expansion ports 550, and a low speed interface/controller 560 connecting to a low speed bus 570 and a storage device 530. Each of the components 510, 520, 530, 540, 550, and 560, are interconnected using various busses, and may be mounted on a common motherboard or in other manners as appropriate. The processor 510 can process instructions for execution within the computing device 500),
comparing phoneme embeddings corresponding to each phoneme of the input token to phoneme embeddings of each phoneme of each content token Jia, para 0013, determining the respective grapheme token representing the respective word of sequence of words corresponding to the respective phoneme token may include determining that the respective grapheme token includes a corresponding word position embedding that matches the respective word position embedding of the respective phoneme token; [i.e., “grapheme token” as “input token” and “word position embedding of the respective phoneme token” as “content token”]).
Therefore, it would have been obvious for one having ordinary skill in the art before the effective filing date of the claimed invention to incorporate the method of Jia in the method of Kumaran in view of Gupta because this would enable the search engine return one or more search results that the device interprets to generate a response for the user (Jia, para 0036).
Regarding Claim 17, Kumaran in view of Gupta, and further in view of Jia disclose the one or more tangible processor-readable storage media of claim 15, wherein the process further comprises:
generating a phonemic index of a first content token corresponding to the first orthography (Kumaran, 16th page, 2nd para, defining a phonemic index based on the classic Soundex algorithm [10] and by adopting q-gram techniques [7] that has been successfully used in approximate matching of monolingual names to multilingual world ); and
Gupta further teaches:
adding the phonemic index [[for]] of the first content token in the first orthography to an inverted index corresponding to the first orthography in the inverted index database (Gupta, page 682, col 2, 2nd para – page 683, 1st col, 3rd para, Now we describe the experimental set up for evaluating the effectiveness of the proposed method for retrieval in Mixed-Script space. 5.1 Dataset We used the FIRE 2013 shared task collection on Transliterated Search [26] for experiments and training. The dataset comprises of document collection, queryset (Q) and relevance judgments. The collection (D1) contains 62,888 documents containing song title and lyrics in Roman, Devanagari and mixed scripts. Statistics of the document collection is given in Table 2 (a). The Q contains 25 lyrics search queries for Bollywood songs in Roman script with mean query length of 4.5 words. Table 2 (b) lists a few examples of queries from Q.
PNG
media_image1.png
96
348
media_image1.png
Greyscale
The experimental setup is a standard adhoc retrieval setting. The document collection is first indexed to create an inverted index and the index lexicon is used as mining lexicon. Being this a lyrics retrieval set up, the sequential information among the terms is crucial for effectiveness evaluation, e.g. “love me baby” and “baby love me” are completely different songs. In order to capture the word-ordering we consider word 2-grams as a unit for indexing and retrieval. The non-trivial part of mixed-script IR is query-enrichment to handle the challenges described in Sec. 2. In order to enrich the query with equivalents, we find the equivalents of the query terms as described in Section 4.5 and the word 2-gram query is formulated as shown in [13]. We consider a variety of systems to be compared with the proposed method. The query formulation is similar for all the systems including the retrieval settings like inverted index, retrieval model and mining lexicon except the method of finding the equivalents… The problem of finding equivalents is formulated as searching across the views by learning hashing functions as presented in [20]… An inverted index of hash codes is prepared for terms in mining lexicon. The equivalents for the query term are found from this index according to the score given by the graph matching algorithm (according to the cosine similarity in the common geometric space) of [30] ).
Regarding Claim 20, Kumaran in view of Gupta, and further in view of Jia disclose the one or more tangible processor-readable storage media of claim 15, wherein the process further comprises:
generating the phonemic index of a phonemic variant of a first content token corresponding to the second orthography (Kumaran, 16th page, 2nd para, defining a phonemic index based on the classic Soundex algorithm [10] and by adopting q-gram techniques [7] that has been successfully used in approximate matching of monolingual names to multilingual world ); and
Gupta further teaches:
adding the phonemic index of the phonemic variant the first content token corresponding to the second orthography to an inverted index corresponding to the second orthography in the inverted index database (Gupta, page 682, col 2, 2nd para – page 683, 1st col, 3rd para, Now we describe the experimental set up for evaluating the effectiveness of the proposed method for retrieval in Mixed-Script space. 5.1 Dataset We used the FIRE 2013 shared task collection on Transliterated Search [26] for experiments and training. The dataset comprises of document collection, queryset (Q) and relevance judgments. The collection (D1) contains 62,888 documents containing song title and lyrics in Roman, Devanagari and mixed scripts. Statistics of the document collection is given in Table 2 (a). The Q contains 25 lyrics search queries for Bollywood songs in Roman script with mean query length of 4.5 words. Table 2 (b) lists a few examples of queries from Q.
PNG
media_image1.png
96
348
media_image1.png
Greyscale
The experimental setup is a standard adhoc retrieval setting. The document collection is first indexed to create an inverted index and the index lexicon is used as mining lexicon. Being this a lyrics retrieval set up, the sequential information among the terms is crucial for effectiveness evaluation, e.g. “love me baby” and “baby love me” are completely different songs. In order to capture the word-ordering we consider word 2-grams as a unit for indexing and retrieval. The non-trivial part of mixed-script IR is query-enrichment to handle the challenges described in Sec. 2. In order to enrich the query with equivalents, we find the equivalents of the query terms as described in Section 4.5 and the word 2-gram query is formulated as shown in [13]. We consider a variety of systems to be compared with the proposed method. The query formulation is similar for all the systems including the retrieval settings like inverted index, retrieval model and mining lexicon except the method of finding the equivalents… The problem of finding equivalents is formulated as searching across the views by learning hashing functions as presented in [20]… An inverted index of hash codes is prepared for terms in mining lexicon. The equivalents for the query term are found from this index according to the score given by the graph matching algorithm (according to the cosine similarity in the common geometric space) of [30]).
Regarding Claim 21, Kumaran in view of Gupta, and further in view of Jia disclose the method of claim 1.
Furthermore, Jia teaches:
wherein the phonemic index of the input token includes phoneme embeddings corresponding to each phoneme of the input token (Jia, para 0006, each token of the plurality of tokens of the input encoder embedding represents a combination of one of a grapheme token embedding or a phoneme token embedding, a segment embedding, a word position embedding, and/or a position embedding. In these examples, identifying the respective word of the sequence of words corresponding to the respective phoneme token may include identifying the respective word of the sequence of words corresponding to the respective phoneme token based on a respective word position embedding associated with the respective phoneme token. Here, determining the respective grapheme token representing the respective word of sequence of words corresponding to the respective phoneme token may include determining that the respective grapheme token includes a corresponding word position embedding that matches the respective word position embedding of the respective phoneme token).
Therefore, it would have been obvious for one having ordinary skill in the art before the effective filing date of the claimed invention to incorporate the method of Jia in the method of Elisha in view of Gupta because this would enable the search engine return one or more search results that the device interprets to generate a response for the user (Jia, para 0036).
Claims 2, 9 and 16 are rejected under 35 U.S.C. 103 as being unpatentable over Kumaran in view of Gupta, further in view of Jia, and further in view of Elisha et al. Pat App No. US 20160275945 A1 (Elisha).
Regarding Claim 2, Kumaran in view of Gupta, and further in view of Jia disclose the method of claim 1, wherein executing the approximate matching analysis comprises (Kumaran, page 13, 2nd para, matching operators and the inherently fuzzy nature of such matching ).
Kumaran in view of Gupta and Jia do not specifically disclose generating a score for each content token based on per-phoneme similarity analysis with the input token, wherein the score represents a combination of measurements of similarity between phoneme embeddings corresponding to each phoneme of the input token compared to phoneme embeddings each corresponding phoneme of each content token.
However, Elisha, in the same field of endeavor, discloses generating a score for each content token based on per-phoneme similarity analysis with the input token, wherein the score represents a combination of measurements of similarity between phoneme embeddings corresponding to each phoneme of the input token compared to phoneme embeddings each corresponding phoneme of each content token (Elisha, para 0071, Figure 3, As shown by block 370, a search platform may return a list of documents wherein the list may be sorted or the documents may be ranked. For example, Solr provides ranking of results that may be used. In some embodiments, a phonetic ranking may be implemented such that a ranking of documents found is according to a match of their content with a phoneme. For example, as described, a textual term received from a user as a search key may be converted to a phoneme and the phoneme may be searched. A ranking unit may rank documents found based on a matching of the content in the documents to the phoneme used in the search; Elisha, para 0009, A system and method may include, in a set of search phoneme strings, at least one phoneme string based on a pre-configured distance from the textual search term. A system and method may statistically calculate a probability of a recognition error for a phoneme; and based on relating a fuzziness parameter value to the probability, select to include or exclude a phonetic transcription in a result of searching for an element in the set of phonetic transcriptions. A system and method may statistically calculate a probability of a recognition error for a phoneme; and based on relating a fuzziness parameter value to the probability, select to include the phoneme in the set of search phoneme strings).
Therefore, it would have been obvious for one having ordinary skill in the art before the effective filing date of the claimed invention to incorporate the method of Elisha in the method of Kumaran in view of Gupta and Jia because this would enable a fast search over large amounts of textual data (Elisha, para 0003).
Regarding Claim 9, Kumaran in view of Gupta, and further in view of Jia disclose the computing system of claim 8, wherein the approximate match analyzer (Kumaran, page 13, 2nd para, matching operators and the inherently fuzzy nature of such matching ).
Jia further discloses to compare phoneme embeddings corresponding to each phoneme for the input token to phoneme embeddings for each phoneme of each content token (Jia, para 0013, determining the respective grapheme token representing the respective word of sequence of words corresponding to the respective phoneme token may include determining that the respective grapheme token includes a corresponding word position embedding that matches the respective word position embedding of the respective phoneme token; [i.e., “grapheme token” as “input token” and “word position embedding of the respective phoneme token” as “content token”]).
Kumaran in view of Gupta and Jia do not specifically disclose generate a score of each content token based on per-phoneme similarity analysis with the input token, wherein the score represents a combination of measurements of similarity between phoneme embeddings corresponding to each phoneme of the input token compared to phoneme embeddings of each corresponding phoneme of each content token.
However, Elisha, in the same field of endeavor, discloses generate a score of each content token based on per-phoneme similarity analysis with the input token, wherein the score represents a combination of measurements of similarity between phoneme embeddings corresponding to each phoneme of the input token compared to phoneme embeddings of each corresponding phoneme of each content token (Elisha, para 0071, Figure 3, As shown by block 370, a search platform may return a list of documents wherein the list may be sorted or the documents may be ranked. For example, Solr provides ranking of results that may be used. In some embodiments, a phonetic ranking may be implemented such that a ranking of documents found is according to a match of their content with a phoneme. For example, as described, a textual term received from a user as a search key may be converted to a phoneme and the phoneme may be searched. A ranking unit may rank documents found based on a matching of the content in the documents to the phoneme used in the search; Elisha, para 0009, A system and method may include, in a set of search phoneme strings, at least one phoneme string based on a pre-configured distance from the textual search term. A system and method may statistically calculate a probability of a recognition error for a phoneme; and based on relating a fuzziness parameter value to the probability, select to include or exclude a phonetic transcription in a result of searching for an element in the set of phonetic transcriptions. A system and method may statistically calculate a probability of a recognition error for a phoneme; and based on relating a fuzziness parameter value to the probability, select to include the phoneme in the set of search phoneme strings).
Therefore, it would have been obvious for one having ordinary skill in the art before the effective filing date of the claimed invention to incorporate the method of Elisha in the method of Kumaran in view of Gupta and Jia because this would enable a fast search over large amounts of textual data (Elisha, para 0003).
Regarding Claim 16, Kumaran in view of Gupta, and further in view of Jia disclose the one or more tangible processor-readable storage media of claim 15 wherein executing the approximate matching analysis (Kumaran, page 13, 2nd para, matching operators and the inherently fuzzy nature of such matching ).
Jia further discloses comparing phoneme embeddings corresponding to each phoneme for the input token to phoneme embeddings for each phoneme of each content token (Jia, para 0013, determining the respective grapheme token representing the respective word of sequence of words corresponding to the respective phoneme token may include determining that the respective grapheme token includes a corresponding word position embedding that matches the respective word position embedding of the respective phoneme token; [i.e., “grapheme token” as “input token” and “word position embedding of the respective phoneme token” as “content token”]).
Kumaran in view of Gupta and Jia do not specifically disclose generating a score for each content token based on per-phoneme similarity analysis with the input token, wherein the score represents a combination of measurements of similarity between phoneme embeddings corresponding to each phoneme of the input token compared to phoneme embeddings of each corresponding phoneme of each content token.
However, Elisha, in the same field of endeavor, discloses generating a score for each content token based on per-phoneme similarity analysis with the input token, wherein the score represents a combination of measurements of similarity between phoneme embeddings corresponding to each phoneme of the input token compared to phoneme embeddings of each corresponding phoneme of each content token (Elisha, para 0071, Figure 3, As shown by block 370, a search platform may return a list of documents wherein the list may be sorted or the documents may be ranked. For example, Solr provides ranking of results that may be used. In some embodiments, a phonetic ranking may be implemented such that a ranking of documents found is according to a match of their content with a phoneme. For example, as described, a textual term received from a user as a search key may be converted to a phoneme and the phoneme may be searched. A ranking unit may rank documents found based on a matching of the content in the documents to the phoneme used in the search; Elisha, para 0009, A system and method may include, in a set of search phoneme strings, at least one phoneme string based on a pre-configured distance from the textual search term. A system and method may statistically calculate a probability of a recognition error for a phoneme; and based on relating a fuzziness parameter value to the probability, select to include or exclude a phonetic transcription in a result of searching for an element in the set of phonetic transcriptions. A system and method may statistically calculate a probability of a recognition error for a phoneme; and based on relating a fuzziness parameter value to the probability, select to include the phoneme in the set of search phoneme strings).
Therefore, it would have been obvious for one having ordinary skill in the art before the effective filing date of the claimed invention to incorporate the method of Elisha in the method of Kumaran in view of Gupta and Jia because this would enable a fast search over large amounts of textual data (Elisha, para 0003).
Claims 4, 11 and 18 are rejected under 35 U.S.C. 103 as being unpatentable over Kumaran in view of Gupta, and further in view of Jia, and further in view of Robertson Pat App No. US 20110106792 A1 (Robertson).
Regarding Claim 4, Kumaran in view of Gupta, and further in view of Jia disclose the method of claim
Gupta further discloses:
generating a phonemic variant of a first content token corresponding to the second orthography (Gupta, page 681, 4th para, Phonemes of the language can be captured by the character n-grams. Consider the feature set F = {f1, . . . , fK} containing character grams of scripts si for all i ∈ {1, ., r} and |F| = K…; Gupta, page 682, 2nd para, the size of phonemes (captured by the character n-grams) for a language is finite and fairly small (only for languages with finite set of alphabets e.g. English with alphabets a-z). Hence, enough evidence for all the phonemes is found even in a small to moderate size training data, which increases the suitability of our approach to the problem; Gupta, page 685, 3rd para, A similar method that uses both stemming and grapheme-to-phoneme conversion is used by [25] to develop a proof-of-concept for a multilingual search engine for 10 Indian languages).
Elisha in view of Gupta and Jia does not specifically disclose using a neural phonemic translation machine learning model trained using phonemic index pairs corresponding to the first language and the second language.
However, Robertson, in the same field of endeavor, discloses using a neural phonemic translation machine learning model trained using phonemic index pairs corresponding to the first language and the second language (Robertson, para 0036-0046, In a preferred embodiment, an approach developed by Kevin Lenzo and Vincent Pagel at Carnegie Melon University (CMU) department of linguistics is used, which uses a phonetic dictionary for direct look-up of phonetic transcriptions whenever possible, but falls back on a set of transcription rules to handle out of dictionary cases. The CMU pronouncing dictionary is publicly available and downloadable from the Internet and is a machine readable pronunciation dictionary for North American English that contains 125,000 words and names together with their phonetic transcriptions in the phoneme set shown in FIG. 1. For words that are not in this dictionary, phonetic transcription is handled by using the large number of transcriptions provided within the dictionary to train a decision tree that is capable of making an accurate determination of the likely pronunciation of any unincluded words. The process for training a decision tree in this manner is described in detail in Pagel V., Lenzo K., and Black A. W. (1998) "Letter to sound rules for accented lexicon compression." Proc. ICSLP, Sydney, Australia, the contents of which are incorporated herein by reference. The CMU pronouncing dictionary comes with a number of Perl scripts which can be used for constructing text-to-phoneme (TTP) decision trees using the Iterative Dichotomiser 3 (ID3) tree learning algorithm... In a preferred embodiment, once the decision tree has been constructed, each word from the original phonetic dictionary is run through the decision tree and compared with the predicted pronunciation within the dictionary. Any words that are correctly predicted by the tree can be eliminated from the dictionary, resulting in a smaller dictionary containing only those words transcribed incorrectly...FIG. 2 illustrates the CMU Dictionary 20, which is used to train the transcription decision tree 21. Once the decision tree is finalised, the words in the CMU Dictionary are passed through the decision tree. Those words which the decision tree correctly transcribes are eliminated from the reduced dictionary. Only those words in the CMU that are not correctly transcribed are retained to form the reduced dictionary 22. In use for transcription, a word 23 is input and the reduced dictionary 22 is first checked. If the input word is in the reduced dictionary, the corresponding phoneme string 24 is output. If the input word is not in the reduced dictionary, the decision tree 21 is then used to transcribe the input word into a string of phonemes 24. The original CMU pronouncing dictionary is 3,507 kb in size, whereas the corresponding decision tree and reduced dictionary occupy just 723 kb and 546 kb respectively. Eliminating words from the dictionary that are correctly transcribed by the decision tree therefore saves considerable memory resources. However, as an alternative, phonetic transcription can be carried out simply by using a phonetic dictionary such as the CMU dictionary. As a further alternative, phonetic transcription can be carried out solely using a decision tree, without reference to any form of dictionary. The choice of dictionary employed for lookup or decision tree training is typically determined by the language being used. Many machine-readable phonetic dictionaries are freely available for download on the internet. Alternatives include BOMP (German), Lexique (French) and the MBRDICO project which provides dictionary resources for several languages. The Unicode Unihan database provides detailed properties and pronunciations for the characters used in Chinese, Japanese and Korean orthography…In order to support phonetic indexing and retrieval of words and names from a database, the present invention incorporates a system for the generation of phonetic index keys. These keys are based on the phoneme transcriptions described above with reference to FIG. 2).
Therefore, it would have been obvious for one having ordinary skill in the art before the effective filing date of the claimed invention to incorporate the method of Robertson in the method of Elisha in view of Gupta and Jia because this would enable this invention provides a system and method for ranking or scoring the degree of similarity between two words based on a comparison of phonemes (Robertson, para 0002).
Regarding Claim 11, Kumaran in view of Gupta, and further in view of Jia disclose the computing system of claim 8.
Gupta further discloses:
to generate a phonemic variant of a first content token corresponding to the second orthography (Gupta, page 681, 4th para, Phonemes of the language can be captured by the character n-grams. Consider the feature set F = {f1, . . . , fK} containing character grams of scripts si for all i ∈ {1, ., r} and |F| = K…; Gupta, page 682, 2nd para, the size of phonemes (captured by the character n-grams) for a language is finite and fairly small (only for languages with finite set of alphabets e.g. English with alphabets a-z). Hence, enough evidence for all the phonemes is found even in a small to moderate size training data, which increases the suitability of our approach to the problem; Gupta, page 685, 3rd para, A similar method that uses both stemming and grapheme-to-phoneme conversion is used by [25] to develop a proof-of-concept for a multilingual search engine for 10 Indian languages).
Elisha in view of Gupta and Jia does not specifically disclose a neural phonemic translation machine learning model executable by the one or more hardware processors and configured …wherein the neural phonemic translation machine learning model is trained using phonemic index pairs corresponding to the first language and the second language.
However, Robertson, in the same field of endeavor, discloses a neural phonemic translation machine learning model executable by the one or more hardware processors and configured …wherein the neural phonemic translation machine learning model is trained using phonemic index pairs corresponding to the first language and the second language ( Robertson, para 0055, These elements are implemented as software modules running on one or more computer processors that are coupled to the name database 40 and key index 41; Robertson, para 0036-0046, In a preferred embodiment, an approach developed by Kevin Lenzo and Vincent Pagel at Carnegie Melon University (CMU) department of linguistics is used, which uses a phonetic dictionary for direct look-up of phonetic transcriptions whenever possible, but falls back on a set of transcription rules to handle out of dictionary cases. The CMU pronouncing dictionary is publicly available and downloadable from the Internet and is a machine readable pronunciation dictionary for North American English that contains 125,000 words and names together with their phonetic transcriptions in the phoneme set shown in FIG. 1. For words that are not in this dictionary, phonetic transcription is handled by using the large number of transcriptions provided within the dictionary to train a decision tree that is capable of making an accurate determination of the likely pronunciation of any unincluded words. The process for training a decision tree in this manner is described in detail in Pagel V., Lenzo K., and Black A. W. (1998) "Letter to sound rules for accented lexicon compression." Proc. ICSLP, Sydney, Australia, the contents of which are incorporated herein by reference. The CMU pronouncing dictionary comes with a number of Perl scripts which can be used for constructing text-to-phoneme (TTP) decision trees using the Iterative Dichotomiser 3 (ID3) tree learning algorithm... In a preferred embodiment, once the decision tree has been constructed, each word from the original phonetic dictionary is run through the decision tree and compared with the predicted pronunciation within the dictionary. Any words that are correctly predicted by the tree can be eliminated from the dictionary, resulting in a smaller dictionary containing only those words transcribed incorrectly...FIG. 2 illustrates the CMU Dictionary 20, which is used to train the transcription decision tree 21. Once the decision tree is finalised, the words in the CMU Dictionary are passed through the decision tree. Those words which the decision tree correctly transcribes are eliminated from the reduced dictionary. Only those words in the CMU that are not correctly transcribed are retained to form the reduced dictionary 22. In use for transcription, a word 23 is input and the reduced dictionary 22 is first checked. If the input word is in the reduced dictionary, the corresponding phoneme string 24 is output. If the input word is not in the reduced dictionary, the decision tree 21 is then used to transcribe the input word into a string of phonemes 24. The original CMU pronouncing dictionary is 3,507 kb in size, whereas the corresponding decision tree and reduced dictionary occupy just 723 kb and 546 kb respectively. Eliminating words from the dictionary that are correctly transcribed by the decision tree therefore saves considerable memory resources. However, as an alternative, phonetic transcription can be carried out simply by using a phonetic dictionary such as the CMU dictionary. As a further alternative, phonetic transcription can be carried out solely using a decision tree, without reference to any form of dictionary. The choice of dictionary employed for lookup or decision tree training is typically determined by the language being used. Many machine-readable phonetic dictionaries are freely available for download on the internet. Alternatives include BOMP (German), Lexique (French) and the MBRDICO project which provides dictionary resources for several languages. The Unicode Unihan database provides detailed properties and pronunciations for the characters used in Chinese, Japanese and Korean orthography…In order to support phonetic indexing and retrieval of words and names from a database, the present invention incorporates a system for the generation of phonetic index keys. These keys are based on the phoneme transcriptions described above with reference to FIG. 2).
Therefore, it would have been obvious for one having ordinary skill in the art before the effective filing date of the claimed invention to incorporate the method of Robertson in the method of Elisha in view of Gupta and Jia because this would enable this invention provides a system and method for ranking or scoring the degree of similarity between two words based on a comparison of phonemes (Robertson, para 0002).
Regarding Claim 18, Kumaran in view of Gupta, and further in view of Jia disclose the one or more tangible processor-readable storage media of claim 15.
Gupta further discloses:
generating a phonemic variant of a first content token corresponding to the second orthography (Gupta, page 681, 4th para, Phonemes of the language can be captured by the character n-grams. Consider the feature set F = {f1, . . . , fK} containing character grams of scripts si for all i ∈ {1, ., r} and |F| = K…; Gupta, page 682, 2nd para, the size of phonemes (captured by the character n-grams) for a language is finite and fairly small (only for languages with finite set of alphabets e.g. English with alphabets a-z). Hence, enough evidence for all the phonemes is found even in a small to moderate size training data, which increases the suitability of our approach to the problem; Gupta, page 685, 3rd para, A similar method that uses both stemming and grapheme-to-phoneme conversion is used by [25] to develop a proof-of-concept for a multilingual search engine for 10 Indian languages).
Elisha in view of Gupta and Jia does not specifically disclose using a neural phonemic translation machine learning model trained using phonemic index pairs corresponding to the first language and the second language.
However, Robertson, in the same field of endeavor, discloses using a neural phonemic translation machine learning model trained using phonemic index pairs corresponding to the original phonetic dictionary is run through the decision tree and compared with the predicted pronunciation within the dictionary. Any words that are correctly predicted by the tree can be eliminated from the dictionary, resulting in a smaller dictionary containing only those words transcribed incorrectly...FIG. 2 illustrates the CMU Dictionary 20, which is used to train the transcription decision tree 21. Once the decision tree is finalised, the words in the CMU Dictionary are passed through the decision tree. Those words which the decision tree correctly transcribes are eliminated from the reduced dictionary. Only those words in the CMU that are not correctly transcribed are retained to form the reduced dictionary 22. In use for transcription, a word 23 is input and the reduced dictionary 22 is first checked. If the input word is in the reduced dictionary, the corresponding phoneme string 24 is output. If the input word is not in the reduced dictionary, the decision tree 21 is then used to transcribe the input word into a string of phonemes 24. The original CMU pronouncing dictionary is 3,507 kb in size, whereas the corresponding decision tree and reduced dictionary occupy just 723 kb and 546 kb respectively. Eliminating words from the dictionary that are correctly transcribed by the decision tree therefore saves considerable memory resources. However, as an alternative, phonetic transcription can be carried out simply by using a phonetic dictionary such as the CMU dictionary. As a further alternative, phonetic transcription can be carried out solely using a decision tree, without reference to any form of dictionary. The choice of dictionary employed for lookup or decision tree training is typically determined by the language being used. Many machine-readable phonetic dictionaries are freely available for download on the internet. Alternatives include BOMP (German), Lexique (French) and the MBRDICO project which provides dictionary resources for several languages. The Unicode Unihan database provides detailed properties and pronunciations for the characters used in Chinese, Japanese and Korean orthography…In order to support phonetic indexing and retrieval of words and names from a database, the present invention incorporates a system for the generation of phonetic index keys. These keys are based on the phoneme transcriptions described above with reference to FIG. 2).
Therefore, it would have been obvious for one having ordinary skill in the art before the effective filing date of the claimed invention to incorporate the method of Robertson in the method of Elisha in view of Gupta and Jia because this would enable this invention provides a system and method for ranking or scoring the degree of similarity between two words based on a comparison of phonemes (Robertson, para 0002).
Claims 5, 12 and 19 are rejected under 35 U.S.C. 103 as being unpatentable over Kumaran in view of Gupta, and further in view of Jia, further in view of Robertson, and further in view of Lawrence Pat App No. GB 2343037 A.
Regarding Claim 5, Kumaran in view of Gupta, and further in view of Jia and Robertson disclose the method of claim 4.
Elisha in view of Gupta, Jia and Robertson do not specifically disclose biasing the phonemic variant of the first content token toward pronunciation of the second orthography.
However, Lawrence, in the same field of endeavor, discloses biasing the phonemic variant of the first content token toward pronunciation of the second orthography (Lawrence, 3rd page, 10th para – 4th page, 3rd para, Using the sorted weightings table the most likely pronunciation is constructed by using the first phonemic variant for each cluster. For “woz” this is "w o z” . A search of the dictionary, using the pronunciation as the key, finds that the orthography pronounced "w o z" is "was". This spelling is saved in the list of possible words presented to the end user. The sorted extract of the weightings table is used to find the next most likely pronunciation. The second entry contains oh "for the second cluster. The pronunciation of "w oh z" is spelt "woes". The third entry contains "E" (schwa). There is no entry in the dictionary for the pronunciation "w E z". When the seventh entry in the table is reached it contains the second phonemic variant for llwll… when the twelfth entry is reached it contains the second phonemic variant for nz". The spelling generator constructs a possible pronunciation using entry 4 for"w", entry 1 for"o"and entry 12 for "z"thus creating w o s". when the twelfth entry is reached it contains the second phonemic variant for nz". The spelling generator constructs a possible pronunciation using entry 4 for"w", entry 1 for"o"and entry 12 for "z"thus creating w o s". There is no entry in the dictionary for nw o sn. The next entries are generated by using"s"for z and all the entries up to entry 12. when entry 14 is reached it contains the third phonemic variant for nWn. The spelling generator uses"v"for the pronunciation of nwn and then uses each variant for the second and third clusters in the order they appear in the table; [“weightings” as “biases”]).
Therefore, it would have been obvious for one having ordinary skill in the art before the effective filing date of the claimed invention to incorporate the method of Lawrence in the method of Elisha in view of Gupta, Jia and Robertson because this would enable searching for possible spellings in the dictionary such that the spellings from the pronunciations with the greatest (heavier) weightings are selected before those with lesser (lighter) weightings (Lawrence, 4th page, 4th para).
Regarding Claim 12, Kumaran in view of Gupta, and further in view of Jia and Robertson disclose the computing system of claim 11.
Elisha in view of Gupta and Jia and Robertson do not specifically disclose wherein the phonemic indexer is further configured to bias the phonemic variant of the first content token toward pronunciation of the second orthography.
However, Lawrence, in the same field of endeavor, discloses wherein the phonemic indexer is further configured to bias the phonemic variant of the first content token toward pronunciation of the second orthography (Lawrence, 3rd page, 10th para – 4th page, 3rd para, Using the sorted weightings table the most likely pronunciation is constructed by using the first phonemic variant for each cluster. For “woz” this is "w o z” . A search of the dictionary, using the pronunciation as the key, finds that the orthography pronounced "w o z" is "was". This spelling is saved in the list of possible words presented to the end user. The sorted extract of the weightings table is used to find the next most likely pronunciation. The second entry contains oh "for the second cluster. The pronunciation of "w oh z" is spelt "woes". The third entry contains "E" (schwa). There is no entry in the dictionary for the pronunciation "w E z". When the seventh entry in the table is reached it contains the second phonemic variant for llwll… when the twelfth entry is reached it contains the second phonemic variant for nz". The spelling generator constructs a possible pronunciation using entry 4 for"w", entry 1 for"o"and entry 12 for "z"thus creating w o s". when the twelfth entry is reached it contains the second phonemic variant for nz". The spelling generator constructs a possible pronunciation using entry 4 for"w", entry 1 for"o"and entry 12 for "z"thus creating w o s". There is no entry in the dictionary for nw o sn. The next entries are generated by using"s"for z and all the entries up to entry 12. when entry 14 is reached it contains the third phonemic variant for nWn. The spelling generator uses"v"for the pronunciation of nwn and then uses each variant for the second and third clusters in the order they appear in the table; [“weightings” as “biases”]).
Therefore, it would have been obvious for one having ordinary skill in the art before the effective filing date of the claimed invention to incorporate the method of Lawrence in the method of Elisha in view of Gupta, Jia and Robertson because this would enable searching for possible spellings in the dictionary such that the spellings from the pronunciations with the greatest (heavier) weightings are selected before those with lesser (lighter) weightings (Lawrence, 4th page, 4th para).
Regarding Claim 19, Kumaran in view of Gupta, and further in view of Jia and Robertson disclose the one or more tangible processor-readable storage media of claim 18.
Elisha in view of Gupta, Jia and Robertson do not specifically disclose biasing the phonemic variant of the first content token toward pronunciation of the second orthography.
However, Lawrence, in the same field of endeavor, discloses biasing the phonemic variant of the first content token toward pronunciation of the second orthography (Lawrence, 3rd page, 10th para – 4th page, 3rd para, Using the sorted weightings table the most likely pronunciation is constructed by using the first phonemic variant for each cluster. For “woz” this is "w o z” . A search of the dictionary, using the pronunciation as the key, finds that the orthography pronounced "w o z" is "was". This spelling is saved in the list of possible words presented to the end user. The sorted extract of the weightings table is used to find the next most likely pronunciation. The second entry contains oh "for the second cluster. The pronunciation of "w oh z" is spelt "woes". The third entry contains "E" (schwa). There is no entry in the dictionary for the pronunciation "w E z". When the seventh entry in the table is reached it contains the second phonemic variant for llwll… when the twelfth entry is reached it contains the second phonemic variant for nz". The spelling generator constructs a possible pronunciation using entry 4 for"w", entry 1 for"o"and entry 12 for "z"thus creating w o s". when the twelfth entry is reached it contains the second phonemic variant for nz". The spelling generator constructs a possible pronunciation using entry 4 for"w", entry 1 for"o"and entry 12 for "z"thus creating w o s". There is no entry in the dictionary for nw o sn. The next entries are generated by using"s"for z and all the entries up to entry 12. when entry 14 is reached it contains the third phonemic variant for nWn. The spelling generator uses"v"for the pronunciation of nwn and then uses each variant for the second and third clusters in the order they appear in the table; [“weightings” as “biases”])).
Therefore, it would have been obvious for one having ordinary skill in the art before the effective filing date of the claimed invention to incorporate the method of Lawrence in the method of Elisha in view of Gupta, Jia and Robertson because this would enable searching for possible spellings in the dictionary such that the spellings from the pronunciations with the greatest (heavier) weightings are selected before those with lesser (lighter) weightings (Lawrence, 4th page, 4th para).
Claim 22 is rejected under 35 U.S.C. 103 as being unpatentable over Kumaran in view of Gupta, and further in view of Jia, further in view of Lee et al. 2021, "Phonetic Variation Modeling and a Language Model Adaptation for Korean English Code-Switching Speech Recognition" Applied Sciences 11, no. 6 (Lee).
Regarding Claim 22, Kumaran in view of Gupta, and further in view of Jia disclose the method of claim 1.
Elisha in view of Gupta, Jia and Robertson do not specifically disclose wherein each phonemic variant is generated by a neural phonemic translation machine learning model trained using phonemic index pairs corresponding to the first language and the second language.
However, Lee, in the same field of endeavor, discloses wherein each phonemic variant is generated by a neural phonemic translation machine learning model trained using phonemic index pairs corresponding to the first language and the second language (Lee, 2nd page, Figure 1. Figure caption, an application of the proposed method: English–Korean automatic speech translator; Lee, 4th page, 1st para, Unlike other languages, Konglish seems to be more severe in its phonetic variations. We should consider the phonetic variations between English pronounced by a Korean who has difficulty speaking English and English pronounced by a Korean who speaks English at a native-like level; Lee, 9th page, 5th para, The base LM was trained based on recurrent neural network (RNN) LM using Korean and English text corpora. The dev. Set consisted of 168 Economy sentences and 1895 lecture sentences. Other parameters and hardware settings are described in Table 4. These are based on the hyperparameters of Librispeech, with some values adjusted; Recently, bidirectional encoder representations from transformers (BERT) [26] represented the semantic relationships of words in an embedding space well… This idea can apply to English–Korean automatic speech translator in Figure 1).
Therefore, it would have been obvious for one having ordinary skill in the art before the effective filing date of the claimed invention to incorporate the method of Lawrence in the method of Kumaran in view of Gupta, Jia and Robertson because this would enable the application of English–Korean automatic speech translator in Figure 1 and the recognition rate between English and Korean improves via the model with the LM domain adaptation (Lee, 8th page, 4th para).
Conclusion
Any inquiry concerning this communication or earlier communications from the examiner should be directed to MULUGETA T. DUGDA whose telephone number is (703)756-1106. The examiner can normally be reached Mon - Fri, 4:30am - 7:00pm.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Paras D. Shah can be reached at 571-270-1650. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/MULUGETA TUJI DUGDA/Examiner, Art Unit 2653
/DOUGLAS GODBOLD/Primary Examiner, Art Unit 2655