DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Information Disclosure Statement
The information disclosure statement (IDS) submitted on 08/01/2023 is being considered by the examiner.
Claim Rejections - 35 USC § 101
35 U.S.C. 101 reads as follows:
Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title.
Claim(s) 1-20 rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more.
Step 1: Independent claims 1, 10, and 17 recite a method, a computer program product, and a system, respectively. These claims therefore invoke a statutory category (process and machine) in Step 1 of the Subject Matter Eligibility Test.
Step 2A, Prong One: Independent claims 1, 10, and 17, under their broadest reasonable interpretation, recite collecting written content, organizing the written content into a corpus, analyzing audio from a video, writing out the audio into a text, comparing the text to written content within the corpus, and adding the text to the corpus. This is an abstract idea in the form of certain methods of organizing human activity (i.e. mental processes such as observation, evaluation, judgement, and opinion). The steps of obtaining data (i.e. written content), organizing the data (i.e. a corpus of written content), analyzing audio and video data, and comparing the audio and video data to the obtained data could be performed by a human using pen and paper or by purely mental reasoning, save for the recitation of generic computer components.
Step 2A, Prong Two: The claims do not integrate the judicial exception into a practical application. The recitation of “a natural language processor”, “a content analyzer”, and “a transcription service” are generic instructions to perform the abstract idea on/using a computer and do not impose a meaningful limit on the judicial exception. The natural language processor, content analyzer, and transcription service are recited at such a high-level of generality and are merely used as a tool to perform the abstract idea faster and more efficiently. The data receiving is a pre-solution activity required to perform the method and does not add a meaningful limitation. Mere data gathering, and analysis do not provide an inventive concept. There is no improvement to the data obtaining, analyzing, or modifying, the functioning of computers, or to any other technology or technical field.
Step 2B: The claims do not include any additional elements that amount to significantly more than the judicial exception. The only additional elements beyond the abstract idea are the natural language processor, the content analyzer, and the transcription service, which perform generic computational functions such as receiving, analyzing, and outputting data. Such elements are well-understood, routine, and conventional within the field.
Accordingly, claims 1, 10, and 17 are directed to an abstract idea and do not include significantly more than the abstract idea itself.
With respect to claims 2, 12, and 18, the claims relate to analyzing the text for additional information, retrieving additional written content from a second source, and adding the additional written content to the corpus. This is a mental process that could be performed by a human using pen and paper or by purely mental reasoning. No additional elements are present.
With respect to claim 3, the claim relates to the user comprising an individual. This is pre-solution activity and therefore does not add a meaningful limitation. No additional elements are present.
With respect to claim 4, the claim relates to the corpus containing technical documentation, articles, and books. This is pre-solution activity and therefore does not add a meaningful limitation. No additional elements are present.
With respect to claims 5 and 13, the claims relate to the corpus comprising sourcing information from an internet source. This is a mental process that could be performed by a human using pen and paper or by purely mental reasoning. No additional elements are present.
With respect to claim 6, the claim relates to updating the corpus comprising adding material from a previous video associated with a user. This is a mental process that could be performed by a human using pen and paper or by purely mental reasoning. No additional elements are present.
With respect to claim 7 and 16, the claims relate to presenting an interface to a user to edit the corpus and to provide feedback on the accuracy of the text. . This is a mental process that could be performed by a human using pen and paper or by purely mental reasoning. The only additional element is “an interface”, which is a generic instruction to perform the abstract idea on/using a computer and does not impose a meaningful limit on the judicial exception. No additional elements are present.
With respect to claims 8, 14, and 19, the claims relate to weighting written content using term frequency-inverse document frequency (TF-IDF). This is a mental process that could be performed by a human using pen and paper or by purely mental reasoning. No additional elements are present.
With respect to claims 9, 15, and 20, the claims relate to weighting written content using cosine similarity. This is a mental process that could be performed by a human using pen and paper or by purely mental reasoning. No additional elements are present.
With respect to claim 11, the claim relates to storing program instructions in a computer readable storage device in a data processing system, and wherein the stored program instructions are transferred over a network from a remote data processing system. This is pre-solution activity and therefore does not add a meaningful limitation. No additional elements are present.
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claim(s) 1-6, 9-13, 15, 17-18, and 20 is/are rejected under 35 U.S.C. 103 as being unpatentable over Lee et al. (US Patent No. 11,308,320), hereinafter referred to as Lee, in view of Tiistola (US Patent No. 12,602,427).
Regarding claim 1, Lee discloses a computer-implemented method comprising: retrieving, using web scraping, written content for a topic from a source (Lee Fig. 4C reference character 601 and "In step 601, the scraper server 121 runs a scraper program to access the external data storage 111. The external data storage 111 may be, for example, a website, database, FTP server, or other data storage," Lee col. 18 lines 53-56);
generating, using a natural language processor, a corpus of reference material for a user using the written content ("In an embodiment, document storage 124 may store the text of a plurality of reference documents. The entire text of each of the reference documents may be stored in the document storage 124. In an embodiment, the text of each of the reference documents may be divided into text segments, where each text segment is stored as an individually indexable element," Lee col. 5 lines 45-51);
and adding to the text taken from the corpus of references (Lee Fig. 6A reference character 606).
However, Lee fails to disclose analyzing, using a content analyzer, an audio of a video for spoken content for a reference in the corpus of reference material; transcribing, using a transcription service, spoken content within an audio of the video into a text; identifying, using the content analyzer, references in the text using a content analyzer, wherein the content analyzer compares the spoken content to written content within the corpus.
Tiistola teaches a system and method for transforming data from streaming media.
Tiistola teaches analyzing, using a content analyzer, an audio of a video for spoken content for a reference in the corpus of reference material (Tiistola Fig. 1 reference character 120 and Fig. 2 reference characters 120 and 204 and "Text-to-speech system 120 is a content transformation system that can transform text content to audio content and audio content to text content. In the context of this description, audio content is not restricted to audio-only content. For example, in some embodiments, audio content can include video content, which can be referred to as multi-media content, and would still be considered audio content for the purposes of this document," Tiistola col. 7 lines 41-48);
transcribing, using a transcription service, spoken content within an audio of the video into a text (Tiistola Fig. 2 reference characters 120 and 204);
identifying, using the content analyzer, references in the text using a content analyzer, wherein the content analyzer compares the spoken content to written content within the corpus (Tiistola Fig. 1 reference character 130 and Fig. 2 reference characters 130 and 206 and "For example, matching and selection system 130 can perform analysis of the output of text-to-speech system 120 or any received content to determine particular characteristics of the content itself, such as a topic or category of content, entities mentioned or suggested by the content, and/or the frequency with which a topic or entity is mentioned, among other characteristics," Tiistola col. 7 lines 61-67 through col. 8 line 1).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified Lee’s disclosure of generating a corpus of reference material through web scraping by including Tiistola’s teaching transcribing audio from video and comparing that to other references within the corpus. This would allow videos and the content they include to be accepted more easily into the corpus. Transcription is a well-known and common technique within the art of audio and language processing, and including this technique to allow the acceptance of video content into the corpus would be an obvious inclusion.
Regarding claim 2, Lee, in view of Tiistola, discloses all of the limitations of claim 1. Lee further discloses further comprising updating the corpus of reference material by: analyzing the text, using the content analyzer, for additional information to be added to the corpus of reference material (Lee Fig. 6A reference characters 603 and 604 shows analyzing the document and its text to then find a similarity to store to the corpus);
retrieving additional written content from a second source based on the additional information in the text (Lee Fig. 6A reference character 605);
and adding the additional written content from a second source based on the additional information to the corpus (Lee Fig. 6A reference character 606).
Regarding claim 3, Lee, in view of Tiistola, discloses all of the limitations of claim 1. Lee further discloses wherein a user comprises an individual (Lee Fig. 11A and 11B show user interfaces, user is inherently understood to be an individual).
Regarding claim 4, Lee, in view of Tiistola, discloses all of the limitations of claim 1. Lee further discloses wherein the corpus of reference material comprises technical documentation, articles, and books ("In some embodiments, the data being scraped may comprise non-patent literature such as scientific articles, magazine articles, websites, images, video, and other reference documents," Lee col. 18 line 67 through col. 19 lines 1-3 and "In one embodiment, the data retrieved by the scraper server 121 also includes the text of documents, such as patents, patent applications, patent publications, scientific articles, magazine articles, and websites," Lee col. 19 lines 9-13).
Regarding claim 5, Lee, in view of Tiistola, discloses all of the limitations of claim 1. Lee further discloses wherein generating the corpus of reference comprises sourcing information from an internet source ("In one embodiment, the data retrieved by the scraper server 121 also includes the text of documents, such as patents, patent applications, patent publications, scientific articles, magazine articles, and websites," Lee col. 19 lines 9-13).
Regarding claim 6, Lee, in view of Tiistola, discloses all of the limitations of claim 2. Lee further discloses wherein updating the corpus of reference comprises adding material from a previous videos associated with a user ("In some embodiments, the data being scraped may comprise non-patent literature such as scientific articles, magazine articles, websites, images, video, and other reference documents," Lee col. 18 line 67 through col. 19 lines 1-3).
Regarding claim 9, Lee, in view of Tiistola, discloses all of the limitations of claim 1. Lee further discloses wherein the written content is weighted using cosine similarity ("The distance between different tensors or encodings may be measured using a distance metric. For example, distance metrics that may be used include cosine similarity, dot product, Euclidean distance, Manhattan distance, Minkowski distance, and others," Lee col. 15 lines 20-24).
As to claim 10, computer program product claim 10 and method claim 1 are related as method and computer program product of using same, with each claimed element’s function corresponding to the method step. Accordingly, claim 10 is similarly rejected under the same rationale as applied above with respect to the method claim.
Regarding claim 11, Lee, in view of Tiistola, discloses all of the limitations of claim 10. Lee further discloses wherein the stored program instructions are stored in a computer readable storage device in a data processing system, and wherein the stored program instructions are transferred over a network from a remote data processing system (Lee Fig. 1 reference character 111 external data storage and reference character 102 network).
As to claims 12-13, computer program product claims 12-13 and method claims 2 and 5 are related as method and computer program product of using same, with each claimed element’s function corresponding to the method step, respectively. Accordingly, claims 12-13 are similarly rejected under the same rationale as applied above with respect to the method claim.
As to claim 15, computer program product claim 15 and method claim 9 are related as method and computer program product of using same, with each claimed element’s function corresponding to the method step. Accordingly, claim 15 is similarly rejected under the same rationale as applied above with respect to the method claim.
As to claims 17-18, system claims 17-18, and method claims 1 and 2 are related as method and system of using same, with each claimed element’s function corresponding to the method step, respectively. Accordingly, claims 17-18 are similarly rejected under the same rationale as applied above with respect to the method claim.
As to claim 20, system claim 20 and method claim 9 are related as method and system of using same, with each claimed element’s function corresponding to the method step. Accordingly, claim 20 is similarly rejected under the same rationale as applied above with respect to the method claim.
Claim(s) 7 and 16 is/are rejected under 35 U.S.C. 103 as being unpatentable over Lee, in view of Tiistola, and further in view of Lewis (US Patent Application Publication No. 2021/0074277).
Regarding claim 7, Lee, in view of Tiistola, discloses all of the limitations of claim 1. Lee further discloses further comprising presenting an interface to a user (Lee Fig. 11A and 11B show user interfaces) wherein the interface allows the user to edit the corpus of reference material (Lee Fig. 11A reference text and parameters can be edited by the user).
However, Lee fails to disclose and to provide feedback on an accuracy of the text.
Lewis teaches a system and method of a transcription revision interface for speech recognition.
Lewis teaches and to provide feedback on an accuracy of the text (Lewis Fig. 2 reference characters 212-216).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified Lee’s disclosure of generating a corpus of reference material through web scraping by including Lewis’s teaching of allowing input/feedback on a transcription. Allowing for a user to correct errors or mistakes on a transcription task is a common and well-known technique within the art of audio and speech processing, as it allows for backchecking to correct mistakes that the system may have made, thereby increasing transcription accuracy as a whole.
As to claim 16, computer program product claim 16 and method claim 7 are related as method and computer program product of using same, with each claimed element’s function corresponding to the method step. Accordingly, claim 16 is similarly rejected under the same rationale as applied above with respect to the method claim.
Claim(s) 8, 14, and 19 is/are rejected under 35 U.S.C. 103 as being unpatentable over Lee, in view of Tiistola, and further in view of Rossillo et al. (US Patent Application Publication No. 2024/0419704), hereinafter referred to as Rossillo.
Regarding claim 8, Lee, in view of Tiistola, discloses all of the limitations of claim 1. However, Lee fails to disclose wherein the written content is weighted using term frequency-inverse document frequency (TF-IDF).
Rossillo teaches a system and method for an adaptive term frequency-inverse document frequency (TF-IDF) inference engine.
Rossillo teaches wherein the written content is weighted using term frequency-inverse document frequency (TF-IDF) (Rossillo Fig. 2 shows an adaptive TF-IDF process).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified Lee’s disclosure of generating a corpus of reference material through web scraping by including Rossillo’s teaching of utilizing TF-IDF as a metric of weighting terms within documents. TF-IDF is a well-known and common technique within text processing and it works to reduce the weight of certain frequent stop words in order to reduce their importance. This would have been an obvious inclusion.
As to claim 14, computer program product claim 14 and method claim 8 are related as method and computer program product of using same, with each claimed element’s function corresponding to the method step. Accordingly, claim 14 is similarly rejected under the same rationale as applied above with respect to the method claim.
As to claim 19, system claim 19 and method claim 8 are related as method and system of using same, with each claimed element’s function corresponding to the method step. Accordingly, claim 19 is similarly rejected under the same rationale as applied above with respect to the method claim.
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure:
US Patent No. 11,315,546
US Patent Application Publication No. 2023/0205985
Any inquiry concerning this communication or earlier communications from the examiner should be directed to ADAM MICHAEL WEAVER whose telephone number is (571)272-7062. The examiner can normally be reached Monday-Friday, 8AM-5PM EST.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Richemond Dorvil can be reached at (571) 272-7602. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/ADAM MICHAEL WEAVER/Examiner, Art Unit 2658
/RICHEMOND DORVIL/Supervisory Patent Examiner, Art Unit 2658