DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Status of the Claims
Claims 1, 8, 15, 19 have been amended and claims 7, 14, 18 have been canceled. Claims 1,3-6,8,10-13, 15-17 and 19-20 are pending.
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claims 1, 3-6, 8, 10-13, 15 is/are rejected under 35 U.S.C. 103 as being unpatentable over Hudetz et al. (US 20240370479) in view of Andreev (US 20150278198) and in further view of Sommer et al. (US 6847966) and Kritt et al. (US 20130311860).
Regarding claim 1, Hudetz teaches a computer-implemented method for searching electronic documents, comprising:
receiving a first natural language semantic search query from a user via a user interface ([0053], [0167]-[0168]) to semantically search a document corpus using a vector index of a plurality of semantically embedded text snippets ([0142]-[0143], [0148]), said vector index facilitating efficient retrieval operations by indexing contextual information for the sequence of words in the text snippets ([0180], [0224]);
servicing the first natural language semantic search query to return a first search result ([0194]), the first search result identifying first documents from the document corpus that are determined to be semantically relevant to the semantic search query ([0148], [0239]-[0240]),
wherein the first search result is presented as a plurality of citations, each citation comprising a chunk of text from a semantically relevant snippet (F17:1706-1712, [0158], [0183]) and a document identifier ([0047] “identify and extract defined sets of information …"information blocks"”, [0122], [0264] “metadata for the electronic document … identification data associated with the electronic document”);
receiving a second lexical search query from the user via the user interface, the second lexical search query comprising search criteria input by the user ([0048] “perform lexical searching, semantic searching, or a combination of both”, [0082], [0131], [0155], [0239])(see NOTE I); and
servicing the second lexical search query to perform a lexical search of the document corpus, wherein the lexical search is constrained to operate only within the first documents identified in the first search result ([0129], [0131], [0147] “drilling down on more specific search questions in follow-up to reviewing previous search results 146”, [0155] “compute a new relevance score over the first set of search results 146 ", [0158] “document corpus returned from the initial result set is analyzed”; “adding an information field with a parameter indicating a type of search, such as "lexical" or "semantic"”, [0131], [0180], [0283] “Additional query generator may be configured to form a query based on … any prior query responses”, [0303] “use determined context of the agreement, its abstractive summary and/or responses to prior queries, to generate a response to the additional query”) to return a second search result ([0155], [0197], [0239]) (see NOTE II) based on a user selection of one or more of the plurality of citations by extracting the document identifier from each of the selected one or more citations ([0255]) and incorporating the extracted document identifier as a parameter in a search request to a second search engine ([0276] “Context information based on … a metadata for the electronic document … one or more identification data associated with the electronic document” [0277] “generate one or more contextualized embeddings for the query. The contextualized embeddings may be configured to form a search vector”, [0290] “search may be performed using one or more identifiers related to the document and/or other information identifying the document”, F23:2302) to return a second search result relevant to the second lexical search query ([0318], [0322], note that context information for additional searches that can include citations [0243], [0255], [0262], F22:2202, 2204) (see NOTE III).
NOTE I Hudetz teaches the present system provides “improved search tools and algorithms to perform lexical searching, semantic searching, or a combination of both”, which “may use the lexical search generator to perform lexical searching in response to a search query” and “may use the semantic search generator to perform semantic searching in response to a search query” and “generate a first set of lexical search results, and … iterate over the first set of lexical search results to generate a second set of semantic search results. Embodiments are not limited in this context” [0131].
Thus, given that lexical and semantic searches are in response to a search query from the user, it would be obvious to one of ordinary skill in the art that a “combination” of such searches can surely allow for various embodiments, such as semantic search to be the first query and a lexical query to be a second query type searching within the semantic results. He show in various embodiments that the semantic query is the first query, which returns “initial result set” (see [0155], [0158]). Given that lexical and semantic searches are performed within the corpus of documents ([0041] “Lexical searching is a process of searching for a particular word or group of words within a given text or corpus”, [0048] “Semantic search … locating the relevant information within an electronic document” also see [0239]), then it is only obvious that the lexical search can be the “second search query from the user to perform a second type of search the document corpus.”
However, to merely obviate such reasoning, Andreev discloses servicing the semantic search query to return a first search result ([0033] “performing semantic searches and clustering search results based on user specified queries”, [0036]) and receiving a second search query from the user to perform a second type of search the document ([0033] “results … clustered by their lexical meanings … allow a user to enter the query and select lexical meanings for one or more words in the selected query … search is performed not only using the words specified in the query, but also the words in specific lexical meanings”, [0072], [0081], [0085]).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the teachings of Hudetz to search semantic results with a different type of query as disclosed by Andreev. Doing so would allow the user to check all variants of meanings of the searched word or word combination and enables searches of not just words or word forms, but of the lexical meaning (Andreev [0001], [0066]).
NOTE II Once again, Hudetz teaches performing “lexical searching, semantic searching, or a combination of both” a single set of search results 146, which can be returned by either lexical or semantic searching – “semantic search generator to iterate over the first set of lexical search results 146 to generate a second set of semantic search results 146”; “use the lexical search … to perform lexical searching in response to a search query 144 … may use the semantic search … to perform semantic searching in response to a search query 144.”
Thus, there is only a single set or search results – 146 and a single query 144. However, a query 144, which can include subsequent queries 144 (see [0197] “a subsequent search query 144”), can be either one - semantic or lexical (defined by the same number 144), which produces either semantic or lexical search results (defined by the same number 146). It is reasonable and obvious to conclude that the subsequent second query 144 (semantic or lexical) which “iterate over the first set of lexical search results 146” can be either one – semantic or lexical, given that the searching can be “a combination of both.” I.e. the search query 144 (which can be semantic or lexical) iterating, searching over search results 146 (which can be semantic or lexical) obviously and reasonably produces all possible combinations of searchings’ – lexical search over the semantic results and a semantic search over the lexical results.
Therefore, it is reasonable and obvious to conclude that given that the semantic query and the lexical query are defined by the same number 144 and a subsequent search query is defined by the same number 144 (which means the subsequent search query 144 can also be semantic or lexical) and the semantic results and the lexical results are defined by the same number 146, produces all possible combinations of searching and is limiting such searching to only search result 146 (which can be semantic or lexical) and satisfies the limitation – “lexical search is constrained to operate only within the first documents identified in the first search result.”
◊ Still, if Hudetz does not explicitly teach, Sommer discloses search is constrained to operate only within the first documents identified in the first search result to return a second search result (C3L54-67 “A term search is accomplished using the Boolean query to search the inverted index to obtain a filtered subset S of matching documents”, C4L3-15, C6L6-7 “comparison is valid only in the defined semantic space for the corpus of information”; C15L12-15 “utilizes only a representative semantic space, document collection 702 may be condensed from c number of document to a smaller number, d, of representative documents for corpus 704”).
It would have been obvious to one of ordinary skill in the art at the time of invention to modify the teachings of He as modified to constrain the second lexical search to the first documents as disclosed by Sommer. Doing so provides an efficient and significant search for achieving overall optimized results and reduces the computational requirements (Sommer C19L31-35, C21L36-37).
NOTE III - Hudetz teaches “perform lexical searching, semantic searching, or a combination of both” [0048] (i.e. first search can be semantic and second search can be lexical) wherein “search may be performed using one or more identifiers related to the document and/or other information identifying the document” [0290]. Sommer further discloses - constraining the search to a representative semantic space condensed to a smaller number (subset) of document (C15L12-15) and narrowing a search further by applying lexical searching on the smaller number (subset) of document defined by the semantic space (C21L12-37, C23L10-40). Thus, it is obvious and reasonable to those skill in the art to combine the teachings of Hudetz and Sommer to perform lexical search within a semantic subset of documents, shown by Sommer, using the document identifiers disclosed by Hudetz to constrain the search, given that Hudetz already using the document identifiers during search to achieve a combination of lexical and semantic searching.
However, to further obviate such reasoning, Kritt discloses - “extraction the document identifier from each of the selected one or more citations and incorporating the extracted document identifier as a parameter in a search request to a second search engine to return a second search result” ([0020]- [0022]). It would have been obvious to one of ordinary skill in the art at the time of invention to modify the teachings of He to search using extracted document identifiers as disclosed by Kritt. Doing so provides a convenient way for the user to explore documents which relate to some topic found within a document located by a search (Kritt [0033]).
Claim 8 recites substantially the same limitations as claim 1, and is rejected for substantially the same reasons.
Regarding claims 3 and 10, Hudetz as modified teaches the method and the medium, further comprises:
providing the first search result to the user in the graphical user interface; and receiving, via user interaction with the graphical user interface, an indication to constrain the second lexical search to the first documents (Hudetz F15-17, [0131], [0147] “drilling down on more specific search questions in follow-up to reviewing previous search results 146”, [0155] “compute a new relevance score over the first set of search results 146 ", [0158] “document corpus returned from the initial result set is analyzed”; “adding an information field with a parameter indicating a type of search, such as "lexical" or "semantic"”, [0131], [0283] “Additional query generator may be configured to form a query based on … any prior query responses”, [0303] “use determined context of the agreement, its abstractive summary and/or responses to prior queries, to generate a response to the additional query”, Andreev [0033], Sommer C12L5-22, C13L18-20, 35-40, C14L21-30, C15L10-19, C16L27-50, C23L10-40).
Regarding claims 4 and 11, Hudetz as modified teaches the method and the medium, further comprising automatically constraining the second lexical search to the first documents (Hudetz F15-17, [0131], Andreev [0033]-[0035], [0079]-[0080], [0084]-[0085], Sommer C12L5-22, C13L18-20, 35-40, C14L21-30, C15L10-19, C16L27-50, C23L10-40).
Regarding claims 5 and 12, Hudetz as modified teaches the method and the medium, wherein servicing the semantic search query comprises:
sending a request to a large language model, the request comprising a prompt to the large language model, the prompt comprising the semantic search query (Hudetz [0137]-[0138], [0159], [0217]-[0236], [0205]);
receiving generative text generated by the large language model in response to the prompt; and including the generative text with the first search result (Hudetz F17:1706-1712, F26-28).
NOTE in analogous art Mukherjee (US 20240354436) likewise teaches claims 5 and 12 in [0046], [0113], [0138], [0146], [0154]-[0155] and further obviates the teaching of Hudetz as modified.
Regarding claims 6 and 13, Hudetz as modified teaches the method and the medium, wherein the request to the large language model includes context to constrain the large language model to the first documents when generating the generative text (Hudetz F13:1308, wherein a “subset” of documents is a constrain, [0155] “compute a new relevance score over the first set of search results 146 ", [0158] “document corpus returned from the initial result set is analyzed”; “adding an information field with a parameter indicating a type of search, such as "lexical" or "semantic"”, [0131], [0283] “Additional query generator may be configured to form a query based on … any prior query responses”, [0303] “use determined context of the agreement, its abstractive summary and/or responses to prior queries, to generate a response to the additional query”).
Regarding claim 15, Hudetz teaches a computer system proving enhanced search, the computer system comprising:
storage storing: a plurality of snippets, each of the plurality of snippets comprising snippet text extracted from a document in a document corpus and a reference to the document from which the snippet text of that snippet was extracted (F17:1706-1712, [0158], [0183]);
an embedding store comprising a vector index of the plurality of semantically embedded snippets, wherein the vector index facilitates efficient retrieval operations by indexing contextual information for a sequence of words in a respective snippet text ([0142]-[0143], [0148], [0180], [0224]); a processor (F2:104);
a semantic search engine executable to perform semantic searching of the document corpus using the vector index perform semantic searching of the document corpus using the vector index to return a first search result identifying first documents from the document corpus that are determined to be semantically relevant to the semantic search query ([0148], [0194], [0239]-[0240]),
wherein the first search result is presented as a plurality of citations, each citation comprising a chunk of text from a semantically relevant snippet (F17:1706-1712, [0158], [0183]);
a lexical search engine executable to perform lexical searching of the document corpus ([0048] “perform lexical searching, semantic searching, or a combination of both”, [0082], [0131], [0155], [0239])(see NOTE I); and
a user interface, wherein the user interface is executable to receive a user selection of one or more of the plurality of citations from the first search result ([0318], [0322], note that context information for additional searches that can include citations [0243], [0255], [0262], F22:2202, 2204) and is further executable to scope lexical searches by the lexical search engine to operate only within the first documents identified by the user selection of the one or more citations ([0131], [0147] “drilling down on more specific search questions in follow-up to reviewing previous search results 146”, [0155] “compute a new relevance score over the first set of search results 146 ", [0158] “document corpus returned from the initial result set is analyzed”; “adding an information field with a parameter indicating a type of search, such as "lexical" or "semantic"”, [0131], [0283] “Additional query generator may be configured to form a query based on … any prior query responses”, [0303] “use determined context of the agreement, its abstractive summary and/or responses to prior queries, to generate a response to the additional query”)(see NOTE II), thereby substantially reducing the amount of data processing and computational resources required for the lexical search engine ([0143] “enabling efficient search and retrieval of information”, [0273] “reduce an amount of processing that may need to be performed”)
NOTE I Hudetz teaches the present system provides “improved search tools and algorithms to perform lexical searching, semantic searching, or a combination of both”, which “may use the lexical search generator to perform lexical searching in response to a search query” and “may use the semantic search generator to perform semantic searching in response to a search query” and “generate a first set of lexical search results, and … iterate over the first set of lexical search results to generate a second set of semantic search results. Embodiments are not limited in this context” [0131].
Thus, given that lexical and semantic searches are in response to a search query from the user, it would be obvious to one of ordinary skill in the art that a “combination” of such searches can surely allow for various embodiments, such as semantic search to be the first query and a lexical query to be a second query type searching within the semantic results. He show in various embodiments that the semantic query is the first query, which returns “initial result set” (see [0155], [0158]). Given that lexical and semantic searches are performed within the corpus of documents ([0041] “Lexical searching is a process of searching for a particular word or group of words within a given text or corpus”, [0048] “Semantic search … locating the relevant information within an electronic document” also see [0239]), then it is only obvious that the lexical search can be the “second search query from the user to perform a second type of search the document corpus.”
However, to merely obviate such reasoning, Andreev discloses servicing the semantic search first ([0033] “performing semantic searches and clustering search results based on user specified queries”, [0036]) and a lexical search engine executable to perform lexical searching of the document corpus as a second search ([0033] “results … clustered by their lexical meanings … allow a user to enter the query and select lexical meanings for one or more words in the selected query … search is performed not only using the words specified in the query, but also the words in specific lexical meanings”, [0072], [0081], [0085]).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the teachings of Hudetz to first perform a semantic search and perform lexical search as a second search as disclosed by Andreev. Doing so would allow the user to check all variants of meanings of the searched word or word combination and enables searches of not just words or word forms, but of the lexical meaning (Andreev [0001], [0066]).
NOTE II Hudetz teaches iterating over the returned search results (which can be lexical or semantic), searching within the previously returned results and generating additional queries based on the returned results, which construed to be analogous to the limitation “scope lexical searches by the lexical search engine to operate only within the first documents.”
Still, if Hudetz does not explicitly teach, Sommer discloses servicing the second lexical search query to perform a lexical search of the document corpus, wherein the lexical search is constrained to operate only within the first documents identified in the first search result to return a second search result (C3L54-67 “A term search is accomplished using the Boolean query to search the inverted index to obtain a filtered subset S of matching documents”, C4L3-15).
Hudetz does not explicitly teach, however Sommer discloses thereby substantially reducing the amount of data processing and computational resources required for the second search engine compared to searching the entire document corpus (C12L5-22, C13L18-20, 35-40, C14L21-30, C15L10-19, C16L27-50, C23L10-40).
It would have been obvious to one of ordinary skill in the art at the time of invention to modify the teachings of He as modified to constrain the second lexical search to the first documents and substantially reducing the amount of data processing as disclosed by Sommer. Doing so provides an efficient and significant search for achieving overall optimized results (Sommer C19L31-35).
Claims 16-20 and alternatively claims 6, 13 is/are rejected under 35 U.S.C. 103 as being unpatentable over Hudetz as modified and in further view of Mukherjee et al. (US 20240354436) based on the provisional application 63/497,932 (dated 04/24/2023).
Regarding claim 16, Hudetz as modified teaches the computer system of Claim 15, wherein:
the semantic search engine is executable to: search the vector index using an embedded query string from a first search input to identify, from the plurality of snippets, semantically relevant snippets that are semantically relevant to the first search input (Hudetz [0142]-[0143], [0148], [0180], [0224]); and
return a corresponding semantic search result to the user interface, the corresponding semantic search result comprising document
the user interface is executable to generate a lexical search request to the lexical search engine to perform a corresponding lexical search, the lexical search request comprising search criteria input by a user (Hudetz [0048] “perform lexical searching, semantic searching, or a combination of both”, [0082], [0131], [0155], [0239]) and the document
Hudetz as modified does not explicitly teach, however Mukherjee discloses search result comprising document identifiers (p.5 13. "Returns the n closest doc ids", p.5 14. "Query the ontology with the returned doc ids to pull the document text", 15). NOTE Mukherjee further discloses search only on the limited set of first documents (p.10 C4).
It would have been obvious to one of ordinary skill in the art at the time of invention to modify the teachings of Hudetz as modified to include document identifiers as disclosed by Mukherjee. Doing so would optimize generating answers from a context (Mukherjee [0008]).
Regarding claim 17, Hudetz as modified teaches the computer system of Claim 16, wherein the user interface is executable to:
display the corresponding semantic search result to the user (Hudetz F17, Mukherjee p.9 C1); and receive, based on a user interaction with the user interface, an indication from the user to scope the corresponding lexical search to the documents identified by the document identifiers from the semantically relevant snippets (Hudetz [0185], [0119], Andreev [0033]-[0035], [0085], Mukherjee p.5, p.9 C1).
Regarding claim 19, Hudetz as modified teaches the computer system of Claim 18, wherein the user interface is executable to:
automatically scope the corresponding lexical search to the documents identified by the document identifiers from the semantically relevant snippets (Hudetz F15-17, [0131], [0147] “drilling down on more specific search questions in follow-up to reviewing previous search results 146”, [0155] “compute a new relevance score over the first set of search results 146 ", [0158] “document corpus returned from the initial result set is analyzed”; “adding an information field with a parameter indicating a type of search, such as "lexical" or "semantic"”, [0131], [0283] “Additional query generator may be configured to form a query based on … any prior query responses”, [0303] “use determined context of the agreement, its abstractive summary and/or responses to prior queries, to generate a response to the additional query”, Sommer C12L5-22, C13L18-20, 35-40, C14L21-30, C15L10-19, C16L27-50, C23L10-40 ,Andreev [0033]-[0035], [0079]-[0080], [0084]-[0085], Mukherjee p.5, p.9 C1).
Regarding claim 20, Hudetz as modified teaches the computer system of Claim 16, wherein the corresponding semantic search result comprises generative text generated by a large language model based on the semantically relevant snippets (Hudetz [0137]-[0138], [0159], [0217]-[0236], [0205], Mukherjee p.5-6, p.9 C1).
Regarding claims 6 and 13, Hudetz as modified teaches the method and the medium, wherein the request to the large language model includes context to constrain the large language model to the first documents when generating the generative text (Hudetz F13:1308, wherein a “subset” of documents is a constrain, [0131], [0147], [0283], [0303]).
Hudetz teaches providing a “a subset of candidate document vectors from the set of candidate document vectors” to the LLM, which obviously constrain the large language model to the first set documents. Alternatively, to further obviate such reasoning Mukherjee teaches request to the large language model includes context to constrain the large language model to the first documents when generating the generative text (Mukherjee p.5-6, p.9 C1). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to modify the teachings of Hudetz to include constrain the large language model to the first documents and include document citations as disclosed by Mukherjee. Doing so would optimize generating answers from a context (Mukherjee [0008]).
◊ Claims 1, 8 and 15 is/are alternatively rejected under 35 U.S.C. 103 as being unpatentable over Hudetz et al. (US 20240370479) in view of Andreev (US 20150278198) and in further view of Sommer et al. (US 6847966) and XU et al. (US 20220121668).
NOTE III - Hudetz teaches “perform lexical searching, semantic searching, or a combination of both” [0048] (i.e. first search can be semantic and second search can be lexical) wherein “search may be performed using one or more identifiers related to the document and/or other information identifying the document” [0290]. Sommer further discloses - constraining the search to a representative semantic space condensed to a smaller number (subset) of document (C15L12-15) and narrowing a search further by applying lexical searching on the smaller number (subset) of document defined by the semantic space (C21L12-37, C23L10-40). Thus, it is obvious and reasonable to those skill in the art to combine the teachings of Hudetz and Sommer to perform lexical search within a semantic subset of documents, shown by Sommer, using the document identifiers disclosed by Hudetz to constrain the search, given that Hudetz already using the document identifiers during search to achieve a combination of lexical and semantic searching.
However, to further obviate such reasoning, XU discloses - “extraction the document identifier from each of the selected one or more citations and incorporating the extracted document identifier as a parameter in a search request to a second search engine to return a second search result” ([0079]). It would have been obvious to one of ordinary skill in the art at the time of invention to modify the teachings of He to search using extracted document identifiers as disclosed by XU. Doing so improves the accuracy of document recommendation and the variety of recommended documents (XU [0045]).
Response to Arguments
Applicant's arguments, filed 05/11/2026, in regard to the presently amended claims have been fully considered and are addressed in the updated rejections to the claims above.
Conclusion
Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a).
A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to POLINA G PEACH whose telephone number is (571)270-7646. The examiner can normally be reached Monday-Friday, 9:30 - 5:30.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Aleksandr Kerzhner can be reached at 571-270-1760. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/POLINA G PEACH/Primary Examiner, Art Unit 2165 May 26, 2026