Prosecution Insights
Last updated: October 02, 2026
Application No. 18/528,116

POSITION-BASED TEXT-TO-SPEECH MODEL

Non-Final OA §101§103
Filed
Dec 04, 2023
Priority
Sep 18, 2023 — provisional 63/583,446
Examiner
HUTCHESON, CODY DOUGLAS
Art Unit
2659
Tech Center
2600 — Communications
Assignee
Adobe Inc.
OA Round
3 (Non-Final)
62%
Grant Probability
Moderate
3-4
OA Rounds
0m
Est. Remaining
99%
With Interview

Examiner Intelligence

Grants 62% of resolved cases
62%
Career Allowance Rate
20 granted / 32 resolved
+0.5% vs TC avg
Strong +38% interview lift
Without
With
+37.5%
Interview Lift
resolved cases with interview
Typical timeline
2y 9m
Avg Prosecution
28 currently pending
Career history
66
Total Applications
across all art units

Statute-Specific Performance

§101
31.4%
-8.6% vs TC avg
§103
45.3%
+5.3% vs TC avg
§102
14.1%
-25.9% vs TC avg
§112
5.4%
-34.6% vs TC avg
Black line = Tech Center average estimate • Based on career data from 32 resolved cases

Office Action

§101 §103
DETAILED ACTION Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Continued Examination Under 37 CFR 1.114 A request for continued examination under 37 CFR 1.114, including the fee set forth in 37 CFR 1.17(e), was filed in this application after final rejection. Since this application is eligible for continued examination under 37 CFR 1.114, and the fee set forth in 37 CFR 1.17(e) has been timely paid, the finality of the previous Office action has been withdrawn pursuant to 37 CFR 1.114. Applicant's submission filed on 04/09/2026 has been entered. Response to Arguments 1. Regarding the rejection under 35 U.S.C. 101, Applicant's arguments filed 04/09/2026 have been fully considered but they are not persuasive. It is argued on pgs. 8-13 of the Remarks that the amended claims are directed to patent-eligible subject matter. Step 2A Prong 1: Applicant argues on pgs. 9-10 that the claims do not recite abstract ideas in the form of mental processes, arguing that the claims recite a specific neural network architecture, multi-task training with simultaneous optimization, spectrogram prediction, and reordered sequence index classification. The Examiner disagrees that the amended claims do not recite abstract ideas. The claims currently recite both mental processes and mathematical concepts, which both fall under the category of abstract ideas. A person can read a document and can write down information about the text they read and the position of this text in the document using pen and paper. A person can further determine speech and a reordered sequence of words to read based on this information. Furthermore, the generation of audio and a reading order via multi-task training and simultaneous optimization amounts to mathematical calculations. The recited machine learning architectures in the claims are considered as additional limitations and thus are considered under Step 2A Prong 2. Therefore, the claims recite abstract ideas under Step 2A Prong 1. Step 2A Prong 2: Applicant argues on pgs. 8-9 and 10-12 that the claims are eligible under Step 2A Prong 2 as integrating any alleged abstract idea into a practical application via technical improvement. Specifically, Applicant argues on pgs. 8-9 that claim 1 recites a specific technical improvement to text-to-speech synthesis technology for semi-structured documents. Applicant further argues on pgs. 10-11 that the claims recite a specific neural network architecture that does not amount to generic computer component implementation. The Examiner respectfully disagrees with these arguments. The claims as currently written do not include sufficient additional limitations to integrate the judicial exception into a practical application. The only additional limitations in claim 1 are a text layout encoder and a reading sequence decoder of an end-to-end text-to-speech model. There are no specific architectures/structure/processing recited for these components which are recited at a high level of generality, and each of these components merely amounts to generic models carrying out the identified abstract ideas. Further details would be needed in the claims in order to fully reflect the technical improvements argued on pgs. 10-11 of the Remarks. Hence, Applicant’s arguments are not persuasive. 2. Regarding the rejections under 35 U.S.C. 103, Applicant’s arguments have been considered but are moot because the new ground of rejection does not rely on any reference applied in the prior rejection of record for any teaching or matter specifically challenged in the argument. Claim Interpretation The following is a quotation of 35 U.S.C. 112(f): (f) Element in Claim for a Combination. – An element in a claim for a combination may be expressed as a means or step for performing a specified function without the recital of structure, material, or acts in support thereof, and such claim shall be construed to cover the corresponding structure, material, or acts described in the specification and equivalents thereof. The following is a quotation of pre-AIA 35 U.S.C. 112, sixth paragraph: An element in a claim for a combination may be expressed as a means or step for performing a specified function without the recital of structure, material, or acts in support thereof, and such claim shall be construed to cover the corresponding structure, material, or acts described in the specification and equivalents thereof. 3. The claims in this application are given their broadest reasonable interpretation using the plain meaning of the claim language in light of the specification as it would be understood by one of ordinary skill in the art. The broadest reasonable interpretation of a claim element (also commonly referred to as a claim limitation) is limited by the description in the specification when 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is invoked. As explained in MPEP § 2181, subsection I, claim limitations that meet the following three-prong test will be interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph: (A) the claim limitation uses the term “means” or “step” or a term used as a substitute for “means” that is a generic placeholder (also called a nonce term or a non-structural term having no specific structural meaning) for performing the claimed function; (B) the term “means” or “step” or the generic placeholder is modified by functional language, typically, but not always linked by the transition word “for” (e.g., “means for”) or another linking word or phrase, such as “configured to” or “so that”; and (C) the term “means” or “step” or the generic placeholder is not modified by sufficient structure, material, or acts for performing the claimed function. Use of the word “means” (or “step”) in a claim with functional language creates a rebuttable presumption that the claim limitation is to be treated in accordance with 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph. The presumption that the claim limitation is interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is rebutted when the claim limitation recites sufficient structure, material, or acts to entirely perform the recited function. Absence of the word “means” (or “step”) in a claim creates a rebuttable presumption that the claim limitation is not to be treated in accordance with 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph. The presumption that the claim limitation is not interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, is rebutted when the claim limitation recites function without reciting sufficient structure, material or acts to entirely perform the recited function. Claim limitations in this application that use the word “means” (or “step”) are being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, except as otherwise indicated in an Office action. Conversely, claim limitations in this application that do not use the word “means” (or “step”) are not being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, except as otherwise indicated in an Office action. This application includes one or more claim limitations that do not use the word “means,” but are nonetheless being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, because the claim limitation(s) uses a generic placeholder that is coupled with functional language without reciting sufficient structure to perform the recited function and the generic placeholder is not preceded by a structural modifier. Such claim limitation(s) is/are: a text-to-phoneme converter module implemented by a machine-learning model to convert text in a digital document into a plurality of phonemes in claim 13 a text-to-speech model implemented by a neural network to convert the plurality of phonemes into digital audio using machine learning in claim 13 a text layout encoder to generate a plurality of text encodings based on the plurality of phonemes using machine learning in claim 13 a reading sequence decoder to decode the plurality of text encodings into the digital audio jointly having a reordered text sequence as part of multi-task learning in claim 13 the reading sequence decoder is configured to generate reordered text sequence in the digital audio which is different from an initial text sequence of the plurality of phonemes in claim 14 the reading sequence decoder is configured to generate the digital audio as including a spectrogram having the reordered text sequence in claim 15 the text layout encoder of the neural network of the text-to-speech model is further configured to generate a text sequence positional encoding as part of the text encoding in claim 17 the reading sequence decoder of the neural network of the text-to-speech model is further configured as a classifier to determine whether a respective said document positional encoding associated with a respective said text encoding indicates a break in the digital document in claim 18 Because this/these claim limitation(s) is/are being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, it/they is/are being interpreted to cover the corresponding structure described in the specification as performing the claimed function, and equivalents thereof. If applicant does not intend to have this/these limitation(s) interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph, applicant may: (1) amend the claim limitation(s) to avoid it/them being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph (e.g., by reciting sufficient structure to perform the claimed function); or (2) present a sufficient showing that the claim limitation(s) recite(s) sufficient structure to perform the claimed function so as to avoid it/them being interpreted under 35 U.S.C. 112(f) or pre-AIA 35 U.S.C. 112, sixth paragraph. Claim Objections 4. Claims 1 and 19 are objected to because of the following informalities: Claim 1: in lines 11-12, “the text-to-speech model implemented using a neural network” should instead be “the text-to-speech model implemented using the neural network”, since antecedent basis is already provided for a neural network in line 6. Claim 19: in line 7, “simultaneous optimization the simultaneous optimization jointly learning…” should instead be “simultaneous optimization, the simultaneous optimization jointly learning…” Appropriate correction is required. Claim Rejections - 35 USC § 101 35 U.S.C. 101 reads as follows: Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title. 5. Claims 1-20 are rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more. Regarding claim 1, “A method” is recited, which is directed to one of the four statutory categories of invention (process) (Step 1: YES). However, the claims limitations, under their broadest reasonable interpretation, recite mental processes or mathematical concepts which fall into the category of abstract idea (Step 2A Prong 1: YES). The following limitations, under their broadest reasonable interpretation, recite mental processes or mathematical concepts: receiving, by a processing device, a digital document having text arranged in an initial text sequence: a person obtains a document having text arranged in an initial sequence generating, by the processing device, a text encoding and a document positional encoding from the digital document…, the document positional encoding is based on a location of the text encoding within the digital document: a person reads the documents, and writes down a text encoding and a positional encoding based on a location of the text within the document, using pen and paper generating, by the processing device, digital audio as part of multi-task training by jointly modeling text reading order detection and digital audio generation…the multi-task training of the neural network of the text-to-speech model jointly learning spectrogram prediction and reordered sequence index classification, as a spectrogram having a reordered text sequence, which is different from the initial text sequence, by decoding the text encoding and the document positional encoding: a person can write down a reordered text sequence different from the initial sequence by decoding the text encoding and positional encoding, using pen and paper. Generating a spectrogram corresponding to the reordered text sequence using multi-task training by jointly modeling text reading order detection and digital audio generation via multi-task training amounts to a mathematical calculation which falls under the abstract idea grouping of mathematical concepts. Claim 1 does not contain any additional elements which integrate the judicial exception into a practical application (Step 2A Prong 2: NO). The only additional limitations are “using a text layout encoder of a text-to-speech model implemented using a neural network”, “using a reading sequence decoder of an end-to-end architecture of the text-to-speech model implemented using a neural network”. These limitations are recited at a high level of generality and amount to mere instructions to implement the judicial exception using a generic computer. Even when viewed in combination with the claim as a whole, mere instructions to implement the judicial exception using a generic computer do not integrate the judicial exception into a practical application as they do not impose any meaningful limits on practicing the abstract idea. Therefore, claim 1 is directed to an abstract idea (Step 2A: YES). Claim 1 does not contain any additional elements which amount to significantly more than the judicial exception (Step 2B: NO). As discussed above, the additional limitations amount to mere instructions to implement the judicial exception using a generic computer which, even when viewed in combination with the claim as a whole, do not integrate the judicial exception into a practical application as they do not impose any meaningful limits on practicing the abstract idea. Therefore, claim 1 is not patent eligible. Regarding dependent claims 2-12, “The method” is recited, which is directed to one of the four statutory categories of invention (process) (Step 1: YES). However, the claims limitations, under their broadest reasonable interpretation, recite mental processes or mathematical concepts which fall into the category of abstract idea (Step 2A Prong 1: YES). The following limitations, under their broadest reasonable interpretation, recite mental processes or mathematical concepts: Claim 2: wherein the document positional encoding is based on coordinates defined in relation to a page of the digital document: a person writes down a positional encoding representing coordinates in relation to the page (e.g. (x,y)) using pen and paper Claim 2 contains no additional limitations. Claim 3: wherein the document positional encoding is based on a bounding box defined for the text: a person writes down a positional encoding representing a bounding box (e.g. writes down coordinates of each corner) using pen and paper Claim 3 contains no additional limitations. Claim 4: wherein the document positional encoding includes four two-dimensional positional encoding defining a relative spatial position of the text within the digital document: a person writes down a positional encoding representing four two-dimensional coordinates in relation to the page (e.g. (x,y)) using pen and paper Claim 4 contains no additional limitations. Claim 5: wherein the generating includes embedding the document positional encoding as part of the text encoding: a person writes down a combined embedding including the positional encoding with the text encoding using pen and paper. Claim 5 contains no additional limitations. Claim 6: generating the text encoding and the document positional encoding…generating the digital audio including the spectrogram having the reordered text sequence…: a person generates text encoding and position encoding, and the reordered text sequence using pen and paper; generating a spectrogram amounts to a mathematical concept. Claim 6 contains the limitations “performed by a text layout encoder of the neural network of the text-to-speech model using machine learning” and “performed using a reading sequence decoder of the neural network of the text-to-speech model using machine learning”. These limitations are recited at a high level of generality and amount to mere instructions to implement the judicial exception using a generic computer. Claim 7: Claim 7 contains the additional limitation “wherein the neural network of the text-to-speech model is trained using curriculum learning”, which amounts to mere instructions to implement the judicial exception using a generic computer. Claim 8: generating the text encoding and the document positional encoding: a person writes down the text and position encodings using pen and paper Claim 8 contains the additional limitation “performed jointly by the text layout encoder”, which amounts to mere instructions to implement the judicial exception using a generic computer. Claim 9: generating the digital audio including the spectrogram having the reordered text sequence: a person writes down a reordered text sequence, and generating a spectrogram amounts to a mathematical concept. Claim 9 contains the additional limitation “performed jointly using the reading sequence decoder”, which amounts to mere instructions to implement the judicial exception using a generic computer. Claim 10: wherein the generating the text encoding further comprises generating a text sequence positional encoding as part of the text encoding…, the text sequence positional encoding defining a position of the text encoding within the text sequence of the digital document: a person writes down a text encoding which comprises a sequence positional encoding defining a position of the text encoding within the text sequence, using pen and paper. Claim 10 contains the additional limitation “by the neural network of the text-to-speech model”, which amounts to mere instructions to implement the judicial exception using a generic computer. Claim 11: wherein the generating includes converting the text from the digital document into a phoneme and wherein the text encoding is generated based on the phoneme: a person reads the text, and converts each word into its corresponding phonemes, and uses the phonemes to generate an encoding using pen and paper. Claim 11 contains no additional limitations. Claim 12: wherein the generating the digital audio includes classifying whether the document positional encoding indicates a break in the digital document: a person determines based on a position encoding if there is a break in the document. Claim 12 contains no additional limitations. Claims 2-12 do not contain any additional elements which integrate the judicial exception into a practical application (Step 2A Prong 2: NO). As discussed above, the only additional limitations are mere instructions to implement the judicial exception using a generic computer which, even when viewed in combination, do not integrate the judicial exception into a practical application because they do not impose any meaningful limits on practicing the abstract idea. Therefore, claims 2-12 are directed to an abstract idea (Step 2A: YES). Claims 2-12 do not contain any additional elements which amount to significantly more than the judicial exception (Step 2B: NO). As discussed above, the only additional limitations are mere instructions to implement the judicial exception using a generic computer, which do not amount to significantly more than the judicial exception as they cannot provide an inventive concept. Therefore, claims 2-12 are not patent eligible. Regarding claim 13, “A system” is recited, which is directed to one of the four statutory categories of invention (machine) (Step 1: YES). However, the claims limitations, under their broadest reasonable interpretation, recite mental processes or mathematical concepts which fall into the category of abstract idea (Step 2A Prong 1: YES). The following limitations, under their broadest reasonable interpretation, recite mental processes or mathematical concepts: convert text in a digital document into a plurality of phonemes: a person reads text in a document and writes down phonemes corresponding to the words using pen and paper. convert the plurality of phonemes into digital audio: a person can use the phonemes to write down audio data using pen and paper generate a plurality of text encodings based on the plurality of phonemes…the plurality of text encoding having embedded, respectively, a document positional encoding based on a location of a respective said text encoding within the digital document: a person a person writes down a text encoding for the phonemes, as well as document positional encoding based on a location of the text within the document, using pen and paper decode the plurality of text encodings into the digital audio jointly having a reordered text sequence as part of multi-task training: a person decodes the text encodings and writes down digital audio corresponding to the encodings using pen and paper. Performing this operation using multi-task training amounts to a mathematical concept. Claim 13 does not contain any additional limitations which integrate the judicial exception into a practical application (Step 2A Prong 2: NO). The only limitations are “a text-to-phoneme converter module implemented by a machine-learning model to…”, “a text-to-speech model implemented by a neural network…using machine learning, the text-to speech model including: a text layer encoder to…using machine learning” and “a reading sequence decoder to…”. These limitations are recited at a high level of generality and amount to mere instructions to implement the judicial exception using a generic computer which, even when viewed in combination, do not integrate the judicial exception into a practical application as they do not impose and meaningful limits on practicing the abstract idea. Therefore, claim 13 is directed to an abstract idea (Step 2A: YES). Claim 13 does not contain any additional elements which amount to significantly more than the judicial exception (Step 2B: NO). As discussed above, the only additional limitations amount to mere instructions to implement the judicial exception using a generic computer, which do not amount to significantly more than the judicial exception as they cannot provide an inventive concept. Therefore, claim 13 is not patent eligible. Regarding dependent claims 14-18, “The system” is recited, which is directed to one of the four statutory categories of invention (machine) (Step 1: YES). However, the claims limitations, under their broadest reasonable interpretation, recite mental processes or mathematical concepts which fall into the category of abstract idea (Step 2A Prong 1: YES). The following limitations, under their broadest reasonable interpretation, recite mental processes or mathematical concepts: Claim 14: generate reordered text sequence in the digital audio which is different from an initial text sequence of the plurality of phonemes: a person writes down a reordered text sequence different from an initial text sequence using pen and paper. Claim 14 contains the additional limitation “wherein the reading sequence decoder is configured to…”. This limitation amounts to mere instructions to implement the judicial exception using a generic computer. Claim 15: generate the digital audio as including a spectrogram having the reordered text sequence: generating a spectrogram amounts to a mathematical concept. Claim 15 contains the additional limitation “wherein the reading sequence is configured to…” which amounts to mere instructions to implement the judicial exception using a generic computer. Claim 16: wherein the document positional encoding is based on coordinates defined in relation to a page of the digital document: a person writes down a positional encoding representing coordinates in relation to the page (e.g. (x,y)) using pen and paper Claim 16 contains no additional limitations. Claim 17: generate a text sequence positional encoding as part of the text encoding, the text sequence positional encoding defining a position of the text encoding within a text sequence of the digital document: a person writes down a positional encoding as part of the text encoding representing a position of the text within the text sequence, using pen and paper. Claim 17 contains the additional limitation “wherein the text layout encoder of the neural network of the text-to-speech model is further configured to…” Claim 18: determine whether a respective said document positional encoding associated with a respective said text encoding indicates a break in the digital document…: a person determines if an encoding corresponding to a break in the digital document. Claim 18 contains the limitations “wherein the reading sequence decoder of the neural network of the text-to-speech model is further configured as a classifier to…”. These limitations are recited at a high level of generality and amount to mere instructions to implement the judicial exception using a generic computer. Claims 14-18 do not contain any additional elements which integrate the judicial exception into a practical application (Step 2A Prong 2: NO). As discussed above, the only additional limitations are mere instructions to implement the judicial exception using a generic computer which, even when viewed in combination, do not integrate the judicial exception into a practical application as they do not impose any meaningful limits on practicing the abstract idea. Therefore, claims 14-18 are directed to an abstract idea (Step 2A: YES). Claims 14-18 do not contain any additional elements which amount to significantly more than the judicial exception (Step 2B: NO). As discussed above, the only additional limitations are mere instructions to implement the judicial exception using a generic computer, which do not amount to significantly more than the judicial exception as they cannot provide an inventive concept. Therefore, claims 14-18 are not patent eligible. Regarding claim 19, “One or more computer readable storage media” is recited, which is directed to one of the four statutory categories of invention (article of manufacture) (Step 1: YES). However, the claims limitations, under their broadest reasonable interpretation, recite mental processes or mathematical concepts which fall into the category of abstract idea (Step 2A Prong 1: YES). The following limitations, under their broadest reasonable interpretation, recite mental processes or mathematical concepts: receive a digital document having text: a person obtains and reads a document having text generating digital audio based on the digital document, the digital audio including a spectrogram having a reading order generated jointly through simultaneous optimization the simultaneous optimization jointly learning spectrogram prediction and reordered sequence index classification: a person writes down audio data based on the document using pen and paper, generating a spectrogram amounts to a mathematical concept. Performing the above using simultaneous optimization amounts to a mathematical calculation. Claim 19 does not contain any additional limitations which integrate the judicial exception into a practical application (Step 2A Prong 2: NO). The only limitations are “One or more computer readable storage media storing instructions that, responsive to execution by a processing device, causes the processing device to perform operations comprising…” and “…generated jointly… by a text layout encoder and a reading order sequence decoder of a neural network of a text-to-speech model”. These limitations are recited at a high level of generality and amount to mere instructions to implement the judicial exception using a generic computer which, even when viewed in combination, do not integrate the judicial exception into a practical application as they do not impose any meaningful limits on practicing the abstract idea. Therefore, claim 19 is directed to an abstract idea (Step 2A: YES). Claim 19 does not contain any additional elements which amount to significantly more than the judicial exception (Step 2B: NO). As discussed above, the only additional limitations amount to mere instructions to implement the judicial exception using a generic computer, which do not amount to significantly more than the judicial exception as they cannot provide an inventive concept. Therefore, claim 19 is not patent eligible. Regarding dependent claim 20, “The one or more computer readable storage media” is recited, which is directed to one of the four statutory categories of invention (article of manufacture) (Step 1: YES). However, the claims limitations, under their broadest reasonable interpretation, recite mental processes or mathematical concepts which fall into the category of abstract idea (Step 2A Prong 1: YES). The following limitations, under their broadest reasonable interpretation, recite mental processes or mathematical concepts: Claim 20: Claim 20 contains the additional limitation “wherein the neural network of the text-to-speech model is trained using curriculum learning”. This limitation amounts to mere instructions to implement the judicial exception using a generic computer. Claim 20 does not contain any additional elements which integrate the judicial exception into a practical application (Step 2A Prong 2: NO). As discussed above, the only additional limitations are mere instructions to implement the judicial exception using a generic computer which, even when viewed in combination, do not integrate the judicial exception into a practical application as they do not impose any meaningful limits on practicing the abstract idea. Therefore, claim 20 is directed to an abstract idea (Step 2A: YES). Claim 20 does not contain any additional elements which amount to significantly more than the judicial exception (Step 2B: NO). As discussed above, the only additional limitations are mere instructions to implement the judicial exception using a generic computer, which do not amount to significantly more than the judicial exception as they cannot provide an inventive concept. Therefore, claim 20 is not patent eligible. Claim Rejections - 35 USC § 103 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. 6. Claims 1, 5-6, 8-9, 11, 13-15, and 19 are rejected under 35 U.S.C. 103 as being unpatentable over Rebecq et al. (US 2024/0257550 A1, hereinafter Rebecq) in view of Abbas et al. (US 11,694,674 B1, hereinafter Abbas) and further in view of Liu et al. (NPL Modeling Prosodic Phrasing with Multi-Task Learning in Tacotron-based TTS, hereinafter Liu). Regarding claim 1, Rebecq discloses A method comprising: receiving, by a processing device (Fig. 5), a digital document having text arranged in an initial text sequence (para. 0055 “As shown in FIG. 4, in step S405 an image representing a document including layout component(s) is received (e.g., received by the encoder 110, for example from a scanning unit, camera or storage medium). For example, the document (e.g., document 105) can include any readable document in the form of an image. The document can include at least one layout component (e.g., paragraphs, summaries, text, images, and/or the like). A reading order for the document can be desired. An example image representing the document can be image 705-1, 705-1-2 shown in FIG. 7.”); generating, by the processing device, a text encoding and a document positional encoding from the digital document using a text layout encoder…implemented using a neural network (para. 0007 “The extracting of the text-based data may include using a neural network configured to generate an embedding including the textual information. The neural network might be a pretrained neural network that maps textual data to an embedding, and an array may include an element including the text-based data associated with each layout component of the plurality of layout components. The identifying of the visual information may include extracting visual-based data from the image. The extracting of the visual-based data may include using a neural network configured to generate an embedding including the visual information. …Also, the textual information might be associated with a first embedding, the visual information might be associated with a second embedding, and the combining of the textual information with the visual information might include concatenating the first embedding with the second embedding.”), the document positional encoding is based on a location of the text encoding within the digital document (para. 0044 “A convolution 325 can be configured to extract features from an image representing the document 105. Features can be based on layout components (e.g., paragraphs, titles, summaries, images, and the like), location of the components, size of the components, color, white space (no components), position of components relative to other components, and/or the like.”; para. 0057 “In step S415 a visual embedding is generated based on the image. For example, the visual embedding can include at least one array including data or visual data (e.g., information and/or features associated with an image) associated with the layout components of the input image (e.g., representing a document). Visual data can include location in the document and relationship to text (e.g., a header associated with the image)”);…modeling text reading order detection (Fig. 3, 310; para. 0050 “The self-attention decoder 355 can be configured to generate a reading order, as ordered layout components 310, based on the context embedding 345...”)…using a reading sequence decoder of an end-to-end architecture… (decoder 355 used to determine reading order 310 in end-to-end architecture Fig. 3; para. 0042 “FIG. 3 illustrates a data flow block diagram according to an example implementation. As shown in FIG. 3, the data flow includes the document 105 (as input), layout components 305, ordered layout components 310, a component model 315 block, and a reading order 320 model block. …Layout components 305 can be strung or linked together with an embedding 330 (e.g., an initial or first embedding) using a layout combiner 335 block which is then input to the reading order model 320 to predict (or generate) the ordered layout components 310.”; Fig. 3, 310; para. 0050 “The self-attention decoder 355 can be configured to generate a reading order, as ordered layout components 310, based on the context embedding 345...”) implemented using a neural network… (para. 0052 “The self-attention decoder 355 can also be composed of a stack of, for example, N=6 identical layers. In addition to the two sub-layers in each encoder layer, the decoder can insert a third sub-layer. The third sub-layer can be configured to perform multi-head attention over the output of the encoder stack. Similar to the encoder, residual connections can be applied around each of the sub-layers, followed by layer normalization…”) learning…reordered sequence index classification…having a reordered text sequence (Fig. 3, model reorders layout components 305 into ordered layout components 310; training to learn reordered sequence index classification: para. 0048 “Each convolution 325 in the component model 315 can have an associated weight. The associated weights can be randomly initialized and then revised in each training iteration (e.g., epoch). The training can be associated with implementing (or helping to implement) distinguishing between layout components and identifying relationships between layout components. In an example implementation, a labeled input image (e.g., document 105 with labels indicating a preferred reading order) and the predicted reading order can be compared. A loss can be generated based on the difference between the labeled reading order and the predicted reading order. Training iterations can continue until the loss is minimized and/or until loss does not change significantly from iteration to iteration. In an example implementation, the lower the loss, the better the predicted reading order.”), which is different from the initial text sequence, by decoding the text encoding and the document positional encoding (Fig. 3, reordered sequence 310 determined via decoding text and document positional encodings: para. 0059 “For example, referring to FIG. 3, the layout combiner 335 can combine the text embedding and the visual embedding.”; combined visual/text embeddings input to 320 to determine reordered sequence: para. 0060 “The self-attention encoder (e.g., self-attention encoder 205, 340) can be configured to generate a context embedding (e.g., context embedding 345).”; para. 0050 “The self-attention decoder 355 can be configured to generate a reading order, as ordered layout components 310, based on the context embedding 345.”). Rebecq does not specifically disclose a text-to-speech model, and generating, by the processing device, digital audio as part of …training by...modeling…digital audio generation…of the text-to-speech model…learning spectrogram prediction… as a spectrogram. Abbas teaches a text-to-speech model (Fig. 2, text-to-speech model), and generating, by the processing device, digital audio as part of …training by...modeling…digital audio generation… (Col. 6 Lines 21-31 “FIG. 4 illustrates embodiments of a text-to-spectrogram (acoustic) model. In some embodiments, this is the acoustic model 122 of FIG. 1. In general, the acoustic model predicts spectrograms at a first level (e.g., sentence-level) and uses that predicted spectrogram at a second level (e.g., word-level). Subsequent levels (e.g., a third level) will use at least the preceding level's predicted spectrogram, but also use all other levels' spectrograms in some embodiments. The predicted spectrograms and upsampled frames of the frame-level are then used to predict an “actual” spectrogram.”; Fig. 4, final spectrogram predicted for input text to 403) of the text-to-speech model…learning spectrogram prediction… as a spectrogram (model is trained to predict spectrograms: Col. 7 Lines 31-35 “In some embodiments, during training, the upsampler 405 is trained by an oracle providing a duration value and/or during inference, the oracle durations are replaced with predicted durations from a durations model 404.”; Col. 7 Lines 45-48 “These vectors are taken in by the word-level frame predictor 417 to generate one or more spectrogram frames (during training, a L1, L2, or other loss function is utilized to train the predictor).”; Col. 7 Lines 60-63 “The phoneme-level frame predictor 409 takes in phoneme embeddings upsampled frames to generate one or more spectrogram frames (during training, a L1, L2, or other loss function is utilized to train the predictor)”). Rebecq and Abbas are considered to be analogous to the claimed invention as they both are in the same field of natural language processing. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have modified Rebecq to incorporate the teachings of Abbas in order to specifically utilize a text-to-speech model to generate digital audio including a spectrogram for the reordered text sequence, and to learn spectrogram prediction. Doing so would be beneficial, as generated spectrograms can be used to produce audio signals (Abbas, Col. 2 Lines 49-57) for the reordered text sequence disclosed in Rebecq, which would lead to more understandable synthesized speech for documents with varieties of digital content in different spatial orientations, improving user experience for those with visual disabilities or who are otherwise busy with other tasks (Tran et al. (US 2021/0020159), Abstract, para. 0001-0002). Rebecq in view of Abbas discloses models trained to learn reordered sequence index classification (see claim mapping above for Rebecq) and spectrogram prediction (see claim mapping above for Abbas). However, Rebecq in view of Abbas does not specifically disclose to [generate the digital audio] as part of multi-task training by jointly [modeling text reading order detection and digital audio generation…] the multi-task training of the neural network of the text-to-speech model jointly [learning spectrogram prediction and reordered sequence index classification…]. Liu teaches to generate the digital audio as part of multi-task training by jointly modeling a main speech generation task and a secondary task (Fig. 1, see “Main task” and “Secondary task” for text to speech model, each associated with a respective loss and used in combination to compute a total loss on the multiple tasks, the model generating output speech (mel-spectrum speech features)), the multi-task training of the neural network of the text-to-speech model jointly learning the speech generation task and the secondary task (Fig. 1, see “Main task” and “Secondary task” for text to speech model, neural network architecture (MTL-Tacotron) trained on the multiple tasks jointly (see Loss_total equation, which combines losses for each respective task); pg. 3, section C “Multi-task Learning…The total loss function is given as Losstotal = Losswav + w*Losspe, with w as a weight…The two-task learning strategy is also referred to as the joint training strategy…”). Rebecq, Abbas, and Liu are considered to be analogous to the claimed invention as they are in the same field of natural language processing. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have modified Rebecq in view of Abbas to incorporate the teachings of Liu in order to specifically generate digital audio as part of multi-task training by jointly modeling text reading order detection and digital audio generation and to jointly learn the disclosed spectrogram prediction and reordered sequence index classification. Doing so would be beneficial, as multi-task learning obtains useful information included in a plurality of different tasks in order to obtain more accurate learning for each individual task, allowing for the different asks to improve each other and prevent a single task from easily falling into a local optimum (Dong et al., US 2025/0061888A1, para. 0047). Regarding claim 5, Rebecq in view of Abbas and Liu discloses wherein the generating includes embedding the document positional encoding as part of the text encoding (Rebecq, para. 0007 “…Also, the textual information might be associated with a first embedding, the visual information might be associated with a second embedding, and the combining of the textual information with the visual information might include concatenating the first embedding with the second embedding.”). Regarding claim 6, Rebecq in view of Abbas and Liu discloses wherein the generating the text encoding and the document positional encoding is performed by a text layout encoder of the neural network of the text-to-speech model using machine learning (Abbas teaches a text-to-speech model (see above claim mapping for Abbas, claim 1); Rebecq discloses generating text and document positional encodings via a text layout encoder: para. 0007 “The extracting of the text-based data may include using a neural network configured to generate an embedding including the textual information. The neural network might be a pretrained neural network that maps textual data to an embedding, and an array may include an element including the text-based data associated with each layout component of the plurality of layout components. The identifying of the visual information may include extracting visual-based data from the image. The extracting of the visual-based data may include using a neural network configured to generate an embedding including the visual information.”); and the generating the digital audio including the spectrogram having the reordered text sequence is performed using a reading sequence decoder of the neural network of the text-to-speech model using machine learning (Rebecq discloses a reading sequence decoder neural network to determine reordered text sequence: Fig. 3, 310; para. 0050 “The self-attention decoder 355 can be configured to generate a reading order, as ordered layout components 310, based on the context embedding 345...”; para. 0052 “The self-attention decoder 355 can also be composed of a stack of, for example, N=6 identical layers. In addition to the two sub-layers in each encoder layer, the decoder can insert a third sub-layer. The third sub-layer can be configured to perform multi-head attention over the output of the encoder stack. Similar to the encoder, residual connections can be applied around each of the sub-layers, followed by layer normalization…”; Abbas teaches generating digital audio including a spectrogram using a text-to-speech model (see above claim mapping for Abbas, claim 1)). Rebecq, Abbas, and Liu are considered to be analogous to the claimed invention as they are in the same field of natural language processing. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have incorporated the teachings of Abbas in order to specifically generate utilize a text-to-speech model and to generate the digital audio including a spectrogram for the reordered text sequence. Doing so would be beneficial, given the same rationale as claim 1. Regarding claim 8, Rebecq in view of Abbas and Liu discloses wherein the generating the text encoding and the document positional encoding is performed jointly by the text layout encoder (Rebecq, para. 0007 “The extracting of the text-based data may include using a neural network configured to generate an embedding including the textual information. The neural network might be a pretrained neural network that maps textual data to an embedding, and an array may include an element including the text-based data associated with each layout component of the plurality of layout components. The identifying of the visual information may include extracting visual-based data from the image. The extracting of the visual-based data may include using a neural network configured to generate an embedding including the visual information. …Also, the textual information might be associated with a first embedding, the visual information might be associated with a second embedding, and the combining of the textual information with the visual information might include concatenating the first embedding with the second embedding.”). Regarding claim 9, Rebecq in view of Abbas and Liu discloses wherein the generating the digital audio including the spectrogram having the reordered text sequence is performed jointly using the reading sequence decoder (Rebecq discloses prediction of reordered text sequence via a reading sequence decoder (Fig. 3, 310, 355); Abbas teaches generation of digital audio including the spectrogram (see above claim mapping for Abbas, claim 1); Liu teaches jointly modeling tasks for text-to-speech processing (see above claim mapping for Liu, claim 1)). Rebecq, Abbas, and Liu are considered to be analogous to the claimed invention as they are in the same field of natural language processing. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have incorporated the teachings of Abbas in order to specifically generate utilize a text-to-speech model and to generate the digital audio including a spectrogram for the reordered text sequence. Doing so would be beneficial, given the same rationale as claim 1. Furthermore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have incorporated the teachings of Liu in order to specifically jointly perform spectrogram prediction and reordered sequence index classification. Doing so would be beneficial, given the same rationale as claim 1. Regarding claim 11, Rebecq in view of Abbas and Liu discloses wherein the generating includes converting the text from the digital document into a phoneme and wherein the text encoding is generated based on the phoneme (Abbas, Col. 7 Lines 9-21 “FIG. 5 illustrates embodiments of an exemplary phoneme embedding encoder. In some embodiments, the phoneme embedding encoder 403 includes one or more of… a tokenizer 503 to tokenize the word(s) of the input text; a grapheme-to-phoneme transcriber 505 to convert graphemes into phonemes (in some embodiments, a set of rules is used to perform this conversion, in other embodiments a neural network is used); an embedding layer INVE07 to embed phonemes into one or more trainable vectors; and a neural network (e.g., BI-LSTM) 509 to generate the phoneme embeddings.”). Rebecq, Abbas, and Liu are considered to be analogous to the claimed invention as they are in the same field of natural language processing. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to incorporate the teachings of Abbas in order to convert the text from the digital document into a phoneme and generating the text encoding based on the phoneme. Doing so would be beneficial, this would allow for phoneme-level spectrograms to be generated which when combined with spectrograms generated for different linguistic units such as sentence and word-level spectrograms, leads to improved naturalness for TTS (Col. 2, Lines 21-32 and 40-48; Fig. 4). Regarding claim 13, Rebecq discloses A system (Fig. 5) comprising:…a text layout encoder to generate a plurality of text encodings…using machine learning (para. 0007 “The extracting of the text-based data may include using a neural network configured to generate an embedding including the textual information. The neural network might be a pretrained neural network that maps textual data to an embedding, and an array may include an element including the text-based data associated with each layout component of the plurality of layout components.), the plurality of text encodings having embedded, respectively, a document positional encoding based on a location of a respective said text encoding within the digital document (para. 0007 “… The identifying of the visual information may include extracting visual-based data from the image. The extracting of the visual-based data may include using a neural network configured to generate an embedding including the visual information. …Also, the textual information might be associated with a first embedding, the visual information might be associated with a second embedding, and the combining of the textual information with the visual information might include concatenating the first embedding with the second embedding.”; para. 0044 “A convolution 325 can be configured to extract features from an image representing the document 105. Features can be based on layout components (e.g., paragraphs, titles, summaries, images, and the like), location of the components, size of the components, color, white space (no components), position of components relative to other components, and/or the like.”; para. 0057 “In step S415 a visual embedding is generated based on the image. For example, the visual embedding can include at least one array including data or visual data (e.g., information and/or features associated with an image) associated with the layout components of the input image (e.g., representing a document). Visual data can include location in the document and relationship to text (e.g., a header associated with the image)”)); and a reading sequence decoder to decode the plurality of text encodings…having a reordered text sequence…( Fig. 3, 335; para. 0050 “The self-attention decoder 355 can be configured to generate a reading order, as ordered layout components 310, based on the context embedding 345...”; para. 0042 “FIG. 3 illustrates a data flow block diagram according to an example implementation. As shown in FIG. 3, the data flow includes the document 105 (as input), layout components 305, ordered layout components 310, a component model 315 block, and a reading order 320 model block. …Layout components 305 can be strung or linked together with an embedding 330 (e.g., an initial or first embedding) using a layout combiner 335 block which is then input to the reading order model 320 to predict (or generate) the ordered layout components 310.”; Fig. 3, 310; para. 0050 “The self-attention decoder 355 can be configured to generate a reading order, as ordered layout components 310, based on the context embedding 345...”). Rebecq does not specifically disclose a text-to-phoneme converter module implemented by a machine-learning model to convert text in a digital document into a plurality of phonemes; and a text-to-speech model implemented by a neural network to convert the plurality of phonemes into digital audio using machine learning, the neural network of the text-to-speech model including…[a reading sequence decode to] decode the plurality of text encodings into the digital audio… Abbas teaches a text-to-phoneme converter module implemented by a machine-learning model to convert text in a digital document into a plurality of phonemes (Col. 7 Lines 9-21 “FIG. 5 illustrates embodiments of an exemplary phoneme embedding encoder. In some embodiments, the phoneme embedding encoder 403 includes one or more of… a tokenizer 503 to tokenize the word(s) of the input text; a grapheme-to-phoneme transcriber 505 to convert graphemes into phonemes (in some embodiments, a set of rules is used to perform this conversion, in other embodiments a neural network is used); an embedding layer INVE07 to embed phonemes into one or more trainable vectors; and a neural network (e.g., BI-LSTM) 509 to generate the phoneme embeddings.”); and a text-to-speech model implemented by a neural network to convert the plurality of phonemes into digital audio using machine learning (Col. 7 Lines 22-24 “The “p” phonemes embeddings are provided to one or more spectrogram generation “levels” within the encoder 401.”; Col. 8 Lines 6-9 “The spectrograms generated in the various levels are concatenated (or otherwise combined) with the one or more frames of the frame-level and fed to the decoder 421 which generates a Mel spectrogram having “t” frames.”), the neural network of the text-to-speech model including…[a reading sequence decode to] decode the plurality of text encodings into the digital audio… (Col. 8 Lines 6-9 “The spectrograms generated in the various levels are concatenated (or otherwise combined) with the one or more frames of the frame-level and fed to the decoder 421 which generates a Mel spectrogram having “t” frames.”). Rebecq and Abbas are considered to be analogous to the claimed invention as they both are in the same field of natural language processing. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have modified Rebecq to incorporate the teachings of Abbas in order to include a text-to-phoneme converter to convert text to a plurality of phonemes, a text-to-speech model to convert the plurality of phonemes into digital audio using machine learning, and to have the reading sequence decoder decode the plurality of text encodings into the digital audio. Using a text-to-speech model to convert phonemes into digital audio would be beneficial, as generated spectrograms can be used to produce audio signals (Abbas, Col. 2 Lines 49-57) for the reordered text sequence disclosed in Cui, which would lead to more understandable synthesized speech for documents with varieties of digital content in different spatial orientations, improving user experience for those with visual disabilities or who are otherwise busy with other tasks (Tran et al. (US 2021/0020159), Abstract, para. 0001-0002). Furthermore, using a text-to-phoneme converter would be beneficial, as this would allow for phoneme-level spectrograms to be generated which when combined with spectrograms generated for different linguistic units such as sentence and word-level spectrograms, leads to improved naturalness for TTS (Col. 2, Lines 21-32 and 40-48; Fig. 4). Rebecq in view of Abbas does not specifically disclose to decode the digital audio jointly having a reordered text sequence as part of multi-task training. Liu teaches to decode the digital audio jointly…as part of multi-task training (Fig. 1, see “Main task” and “Secondary task” for text to speech model, each associated with a respective loss and used in combination to compute a total loss on the multiple tasks, the model generating output speech (mel-spectrum speech features); neural network architecture (MTL-Tacotron) trained on the multiple tasks jointly (see Loss_total equation, which combines losses for each respective task); pg. 3, section C “Multi-task Learning…The total loss function is given as Losstotal = Losswav + w*Losspe, with w as a weight…The two-task learning strategy is also referred to as the joint training strategy…”). Rebecq, Abbas, and Liu are considered to be analogous to the claimed invention as they are in the same field of natural language processing. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have modified Rebecq in view of Abbas to incorporate the teachings of Liu in order to specifically generate digital audio jointly as part of multi-task training. Doing so would be beneficial, as multi-task learning obtains useful information included in a plurality of different tasks in order to obtain more accurate learning for each individual task, allowing for the different asks to improve each other and prevent a single task from easily falling into a local optimum (Dong et al., US 2025/0061888A1, para. 0047). Regarding claim 14, Rebecq in view of Abbas and Liu discloses wherein the reading sequence decoder is configured to generate reordered text sequence in the digital audio which is different from an initial text sequence of the plurality of phonemes (Rebecq discloses a reading sequence decoder to generate a reordered text sequence different from an initial text sequence 305: Fig. 3, reordered sequence 310 determined via decoding text and document positional encodings: para. 0059 “For example, referring to FIG. 3, the layout combiner 335 can combine the text embedding and the visual embedding.”; combined visual/text embeddings input to 320 to determine reordered sequence: para. 0060 “The self-attention encoder (e.g., self-attention encoder 205, 340) can be configured to generate a context embedding (e.g., context embedding 345).”; para. 0050 “The self-attention decoder 355 can be configured to generate a reading order, as ordered layout components 310, based on the context embedding 345.”; Abbas teaches generating digital audio based on an initial text sequence of phonemes: Col. 7 Lines 22-24 “The “p” phonemes embeddings are provided to one or more spectrogram generation “levels” within the encoder 401.”; Col. 8 Lines 6-9 “The spectrograms generated in the various levels are concatenated (or otherwise combined) with the one or more frames of the frame-level and fed to the decoder 421 which generates a Mel spectrogram having “t” frames.”). Rebecq, Abbas, and Liu are considered to be analogous to the claimed invention as they are in the same field of natural language processing. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have incorporated the teachings of Abbas in order to specifically generate digital audio for the reordered text sequence. Doing so would be beneficial, given the same rationale as claim 13. Regarding claim 15, Rebecq in view of Abbas and Liu discloses wherein the reading sequence decoder is configured to generate the digital audio as including a spectrogram having the reordered text sequence (Rebecq discloses a reading sequence decoder configured to generate a reordered text sequence: Fig. 3, reordered sequence 310 determined via decoding text and document positional encodings: para. 0059 “For example, referring to FIG. 3, the layout combiner 335 can combine the text embedding and the visual embedding.”; combined visual/text embeddings input to 320 to determine reordered sequence: para. 0060 “The self-attention encoder (e.g., self-attention encoder 205, 340) can be configured to generate a context embedding (e.g., context embedding 345).”; para. 0050 “The self-attention decoder 355 can be configured to generate a reading order, as ordered layout components 310, based on the context embedding 345.”; Abbas teaches generation of digital audio as including a spectrogram: Col. 7 Lines 22-24 “The “p” phonemes embeddings are provided to one or more spectrogram generation “levels” within the encoder 401.”; Col. 8 Lines 6-9 “The spectrograms generated in the various levels are concatenated (or otherwise combined) with the one or more frames of the frame-level and fed to the decoder 421 which generates a Mel spectrogram having “t” frames.”). Rebecq, Abbas, and Liu are considered to be analogous to the claimed invention as they are in the same field of natural language processing. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have incorporated the teachings of Abbas in order to specifically generate digital audio including a spectrogram for the reordered text sequence. Doing so would be beneficial, given the same rationale as claim 13. Regarding claim 19, Rebecq discloses One or more computer readable storage media storing instructions that, responsive to execution by a processing device, causes the processing device to perform operations including (para. 0005 “In a general aspect, a device, a system, a non-transitory computer-readable medium (having stored thereon computer executable program code which can be executed on a computer system), and/or a method can perform a process with a method including receiving an image representing a document including a plurality of layout components, identifying textual information associated with the plurality of layout components, identifying visual information associated with the plurality of layout components, combining the textual information with the visual information, and predicting a reading order of the plurality of layout components based on the combined textual information and visual information using a self-attention encoder/decoder.”): receiving a digital document having text (para. 0055 “As shown in FIG. 4, in step S405 an image representing a document including layout component(s) is received (e.g., received by the encoder 110, for example from a scanning unit, camera or storage medium). For example, the document (e.g., document 105) can include any readable document in the form of an image. The document can include at least one layout component (e.g., paragraphs, summaries, text, images, and/or the like). A reading order for the document can be desired. An example image representing the document can be image 705-1, 705-1-2 shown in FIG. 7.”); and…a reading order generated (Fig. 3, 310; para. 0050 “The self-attention decoder 355 can be configured to generate a reading order, as ordered layout components 310, based on the context embedding 345...”)…learning…reordered sequence index classification (Fig. 3, model reorders layout components 305 into ordered layout components 310; training to learn reordered sequence index classification: para. 0048 “Each convolution 325 in the component model 315 can have an associated weight. The associated weights can be randomly initialized and then revised in each training iteration (e.g., epoch). The training can be associated with implementing (or helping to implement) distinguishing between layout components and identifying relationships between layout components. In an example implementation, a labeled input image (e.g., document 105 with labels indicating a preferred reading order) and the predicted reading order can be compared. A loss can be generated based on the difference between the labeled reading order and the predicted reading order. Training iterations can continue until the loss is minimized and/or until loss does not change significantly from iteration to iteration. In an example implementation, the lower the loss, the better the predicted reading order.”), by a text layout encoder (para. 0007 “The extracting of the text-based data may include using a neural network configured to generate an embedding including the textual information. The neural network might be a pretrained neural network that maps textual data to an embedding, and an array may include an element including the text-based data associated with each layout component of the plurality of layout components. The identifying of the visual information may include extracting visual-based data from the image. The extracting of the visual-based data may include using a neural network configured to generate an embedding including the visual information. …) and a reading order sequence decoder of a neural network…(Fig. 3, 310; para. 0050 “The self-attention decoder 355 can be configured to generate a reading order, as ordered layout components 310, based on the context embedding 345...”). Rebecq does not specifically disclose generating digital audio based on the digital document, the digital audio including a spectrogram…learning spectrogram prediction...by [a text layout encoder and a reading order sequence decoder of a neural network] of a text-to-speech model. Abbas teaches generating digital audio based on the digital document, the digital audio including a spectrogram… (Col. 6 Lines 21-31 “FIG. 4 illustrates embodiments of a text-to-spectrogram (acoustic) model. In some embodiments, this is the acoustic model 122 of FIG. 1. In general, the acoustic model predicts spectrograms at a first level (e.g., sentence-level) and uses that predicted spectrogram at a second level (e.g., word-level). Subsequent levels (e.g., a third level) will use at least the preceding level's predicted spectrogram, but also use all other levels' spectrograms in some embodiments. The predicted spectrograms and upsampled frames of the frame-level are then used to predict an “actual” spectrogram.”; Fig. 4, final spectrogram predicted for input text to 403) learning spectrogram prediction... (model is trained to predict spectrograms: Col. 7 Lines 31-35 “In some embodiments, during training, the upsampler 405 is trained by an oracle providing a duration value and/or during inference, the oracle durations are replaced with predicted durations from a durations model 404.”; Col. 7 Lines 45-48 “These vectors are taken in by the word-level frame predictor 417 to generate one or more spectrogram frames (during training, a L1, L2, or other loss function is utilized to train the predictor).”; Col. 7 Lines 60-63 “The phoneme-level frame predictor 409 takes in phoneme embeddings upsampled frames to generate one or more spectrogram frames (during training, a L1, L2, or other loss function is utilized to train the predictor)”) by… a text-to-speech model (Fig. 2, text-to-speech model). Rebecq and Abbas are considered to be analogous to the claimed invention as they both are in the same field of natural language processing. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have modified Rebecq to incorporate the teachings of Abbas in order to specifically utilize a text-to-speech model to generate digital audio including a spectrogram for the reordered text sequence, and to learn spectrogram prediction. Doing so would be beneficial, as generated spectrograms can be used to produce audio signals (Abbas, Col. 2 Lines 49-57) for the reordered text sequence disclosed in Rebecq, which would lead to more understandable synthesized speech for documents with varieties of digital content in different spatial orientations, improving user experience for those with visual disabilities or who are otherwise busy with other tasks (Tran et al. (US 2021/0020159), Abstract, para. 0001-0002). Rebecq in view of Abbas discloses models trained to learn reordered sequence index classification (see claim mapping above for Rebecq) and spectrogram prediction (see claim mapping above for Abbas). However, Rebecq in view of Abbas does not specifically disclose [generating digital audio…] jointly through simultaneous optimization the simultaneous optimization jointly learning [spectrogram prediction and reordered sequence index classification]. Liu teaches [generating digital audio…] jointly through simultaneous optimization the simultaneous optimization jointly learning a main speech generation task and a secondary task (Fig. 1, see “Main task” and “Secondary task” for text to speech model, each associated with a respective loss and used in combination to compute a total loss on the multiple tasks, the model generating output speech (mel-spectrum speech features; neural network architecture (MTL-Tacotron) trained on the multiple tasks jointly (see Loss_total equation, which combines losses for each respective task), resulting in simultaneous optimization; pg. 3, section C “Multi-task Learning…The total loss function is given as Losstotal = Losswav + w*Losspe, with w as a weight…The two-task learning strategy is also referred to as the joint training strategy…”). Rebecq, Abbas, and Liu are considered to be analogous to the claimed invention as they are in the same field of natural language processing. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have modified Rebecq in view of Abbas to incorporate the teachings of Liu in order to specifically generate digital audio jointly through simultaneous optimization, the simultaneous optimization jointly learning the spectrogram prediction and reordered sequence index classification. Doing so would be beneficial, as multi-task learning obtains useful information included in a plurality of different tasks in order to obtain more accurate learning for each individual task, allowing for the different asks to improve each other and prevent a single task from easily falling into a local optimum (Dong et al., US 2025/0061888A1, para. 0047). 7. Claims 2-4, 10, and 16-17 are rejected under 35 U.S.C. 103 as being unpatentable over Rebecq in view of Abbas and Liu, and further in view of Cui et al. (US 2024/0265206 A1, hereinafter Cui). Regarding claim 2, Rebecq in view of Abbas and Liu does not specifically disclose wherein the document positional encoding is based on coordinates defined in relation to a page of the digital document. Cui teaches wherein the document positional encoding is based on coordinates defined in relation to a page of the digital document (see Eq. 4; para. 0057 “…the embedding representation 331 of the layout information may be represented as follows… where W and H represent a total width and a total height of the document 162, and L represents a length of the input sequence S corresponding to the text element. In the above Equation (4), the x-axis coordinate information (x0, x1) and width w are used as a triple to construct an embedding representation, and the y-axis coordinate information (y0, y1) and height h are used as a triple to construct another embedding representation, and then the two embedding representations are concatenated into the embedding representation of the layout information of the i.sup.th text element.”). Rebecq, Abbas, Liu, and Cui are considered to be analogous to the claimed invention as they are in the same field of natural language processing. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have modified Rebecq in view of Abbas and Liu to incorporate the teachings of Cui in order to specifically have the document positional encoding be based on coordinates defined in relation to a page of the digital document. Doing so would be beneficial, providing a standardized way of representing relative position within the digital document. Regarding claim 3, Rebecq in view of Abbas and Liu does not specifically disclose wherein the document positional encoding is based on a bounding box defined for the text. Cui teaches wherein the document positional encoding is based on a bounding box defined for the text (para. 0049 “Certainly, in other examples, the layout information of the text element may also be characterized in other ways, for example, the relative spatial position may be represented with a center point of the bounding box, and the size may be represented with an area of the boundary box, and so on. The layout information is not limited in the text herein, as long as it can be ensured that any alternative or additional information of different text elements in the two-dimensional space of the document 162 all may be used. The layout information may also be converted into an embedding representation 331 in the form of a vector for input to the feature extraction model 120.”). Rebecq, Abbas, Liu, and Cui are considered to be analogous to the claimed invention as they are in the same field of natural language processing. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have modified Rebecq in view of Abbas and Liu to incorporate the teachings of Cui in order to specifically have the document positional encoding be based on a bounding box defined for the text. Doing so would be beneficial, providing a standardized way of representing relative position within the digital document. Regarding claim 4, Rebecq in view of Abbas and Liu does not specifically disclose wherein the document positional encoding includes four two-dimensional positional encoding defining a relative spatial position of the text within the digital document. Cui teaches wherein the document positional encoding includes four two-dimensional positional encoding defining a relative spatial position of the text within the digital document (para. 0048 “As an example, the layout information of the i.sup.th text element may be represented as (x.sub.0,x.sub.1,custom-character.sub.0,custom-character.sub.1,w,h), where (x0, y0) represents x-axis and y-axis coordinates of the upper left (right) corner of the bounding box of the text element, (x1, y1) represents x-axis and y-axis coordinates of the lower right (left) corner of the bounding box…”). Rebecq, Abbas, Liu, and Cui are considered to be analogous to the claimed invention as they are in the same field of natural language processing. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have modified Rebecq in view of Abbas and Liu to incorporate the teachings of Cui in order to specifically have the document positional encoding include four two-dimensional positional encoding defining a relative spatial position of the text within the digital document. Doing so would be beneficial, providing a standardized way of representing relative position within the digital document. Regarding claim 10, Rebecq in view of Abbas and Liu discloses the text encoding by the neural network of the text-to-speech model (see above claim mapping for claim 1), but does not specifically disclose wherein the generating the text encoding further comprises generating a text sequence positional encoding as part of the text encoding [by the neural network of the text-to-speech model], the text sequence positional encoding defining a position of the text encoding within a text sequence of the digital document. Cui teaches wherein the generating the text encoding further comprises generating a text sequence positional encoding as part of the text encoding…, the text sequence positional encoding defining a position of the text encoding within a text sequence of the digital document (para. 0050 “In some implementations, the text elements in the text sequence 312 may also have respective sequence index information, which is to indicate sequential positions in the text sequence 312. Different from the two-dimensional relative spatial positions of the document 162 indicated by the layout information, the sequence index information is used to indicate a relative position of the text element in a one-dimensional text sequence 312, and therefore may also be regarded as one-dimensional position information. It is possible to assign corresponding sequence index information to each text element in order from a starting text element of the text sequence 312.”). Rebecq, Abbas, Liu, and Cui are considered to be analogous to the claimed invention as they are in the same field of natural language processing. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have modified Rebecq in view of Abbas and Liu to incorporate the teachings of Cui in order to specifically have the generating the text comprise generating a text sequence positional encoding as part of the text encoding, the text sequence positional encoding defining a position of the text encoding within a text sequence of the digital document. Doing so would be beneficial, providing a standardized way of representing relative position within the digital document. Regarding claim 16, claim 16 is rejected for analogous reasons to claim 2. Regarding claim 17, claim 17 is rejected for analogous reasons to claim 10. 8. Claims 7 and 20 are rejected under 35 U.S.C. 103 as being unpatentable over Rebecq in view of Abbas and Liu, and in further view of Hwang & Chang (NPL Document-Level Neural TTS Using Curriculum Learning and Attention Masking, hereinafter Hwang). Regarding claim 7, Rebecq in view of Abbas and Liu does not specifically disclose wherein the neural network of the text-to-speech model is trained using curriculum learning. Hwang teaches wherein the neural network of the text-to-speech model is trained using curriculum learning (Figure 3; pg. 3 2nd para. “Because our purpose is to synthesize document-level text into speech, we begin training with short sentences and gradually adopt long sentences based on curriculum learning…”). Rebecq, Abbas, Liu, and Hwang are considered to be analogous to the claimed invention as they are all in the same field of natural language processing. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to incorporate the teachings of Hwang in order to train the text-to-speech model using curriculum learning. Doing so would be beneficial, as this would allow for the text-to-speech model to be trained on long sentences with limited GPU capacity (Hwang, Abstract). Regarding claim 20, claim 20 is rejected for analogous reasons to claim 7. 9. Claims 12 and 18 are rejected under 35 U.S.C. 103 as being unpatentable over Rebecq in view of Abbas and Liu, and in further view of Klimkov et al. (NPL Phrase break prediction for long-form reading TTS: exploiting text structure information, hereinafter Klimkov). Regarding claim 12, Rebecq in view of Abbas and Liu does not specifically disclose wherein the generating the digital audio includes classifying whether the document position encoding indicates a break in the digital document. Klimkov teaches wherein the generating the digital audio includes classifying whether the document position encoding indicates a break in the digital document (pg. 2 section 3 “Input Features”: “Distance: CART models do not take context into account. Additional distance features are therefore needed. The number of syllables from the current word to the previous and next punctuation mark was used. This results in 2 additional features…”; pg. 3 section 4 “Modelling”: 2nd para. “The last layer is a softmax which estimates the posterior probabilities of no break and of a respiratory break with a pause…”; pg. 3 section 6 “Subjective evaluations”: 2nd para. “For the listening tests, text with breaks inserted by various phrasing models, was synthesized using a hybrid TTS system…”). Rebecq, Abbas, Liu, and Klimkov are considered to be analogous to the claimed invention as they are all in the same field of natural language processing. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to incorporate the teachings of Klimkov in order to further have the generation of digital audio include classifying whether the document position encoding indicates a break in the digital document. Doing so would be beneficial, as this would improve prediction of phrase breaks for long sentences, increasing the naturalness of the synthesized speech (Klimkov, Abstract). Regarding claim 18, Rebecq in view of Abbas and Liu does not specifically disclose wherein the reading sequence decoder of the neural network of the text-to-speech model is further configured as a classifier to determine whether a respective said document positional encoding associated with a respective said text encoding indicates a break in the digital document. Klimkov teaches wherein the reading sequence decoder is further configured as a classifier to determine whether a respective said document positional encoding associated with a respective said text encoding indicates a break in the digital document (pg. 2 section 3 “Input Features”: “Distance: CART models do not take context into account. Additional distance features are therefore needed. The number of syllables from the current word to the previous and next punctuation mark was used. This results in 2 additional features…”; pg. 3 section 4 “Modelling”: 2nd para. “The last layer is a softmax which estimates the posterior probabilities of no break and of a respiratory break with a pause…”; pg. 3 section 6 “Subjective evaluations”: 2nd para. “For the listening tests, text with breaks inserted by various phrasing models, was synthesized using a hybrid TTS system…”). Rebecq, Abbas, Liu, and Klimkov are considered to be analogous to the claimed invention as they are all in the same field of natural language processing. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to incorporate the teachings of Klimkov in order to further have the generation of digital audio include classifying whether the document position encoding indicates a break in the digital document. Doing so would be beneficial, as this would improve prediction of phrase breaks for long sentences, increasing the naturalness of the synthesized speech (Klimkov, Abstract). Conclusion The prior art made of record and not relied upon is considered pertinent to applicant's disclosure: Agrawal (US 2020/0311185 A1): determination of reading order of objects in digital documents (Fig. 2, Fig. 12) Sodhani et al. (US 2019/0188463 A1): reading order prediction neural network for digital document (Fig. 3a, 3d) Any inquiry concerning this communication or earlier communications from the examiner should be directed to CODY DOUGLAS HUTCHESON whose telephone number is (703)756-1601. The examiner can normally be reached M-F 8:00AM-5:00PM EST. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Pierre-Louis Desir can be reached at (571)-272-7799. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /CODY DOUGLAS HUTCHESON/Examiner, Art Unit 2659 /PIERRE LOUIS DESIR/Supervisory Patent Examiner, Art Unit 2659
Read full office action

Prosecution Timeline

Show 3 earlier events
Nov 24, 2025
Applicant Interview (Telephonic)
Nov 25, 2025
Response Filed
Feb 23, 2026
Final Rejection mailed — §101, §103
Apr 09, 2026
Examiner Interview Summary
Apr 09, 2026
Applicant Interview (Telephonic)
Apr 09, 2026
Request for Continued Examination
Apr 12, 2026
Response after Non-Final Action
Aug 26, 2026
Non-Final Rejection mailed — §101, §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12743450
QUERY FORMATTING SYSTEM, QUERY FORMATTING METHOD, AND INFORMATION STORAGE MEDIUM
3y 6m to grant Granted Sep 22, 2026
Patent 12737563
REGIONAL SIGN LANGUAGE TRANSLATION
3y 5m to grant Granted Sep 15, 2026
Patent 12718828
AUDIO SOURCE CLASSIFICATION FOR HANDSFREE COMMUNICATIONS
3y 5m to grant Granted Aug 25, 2026
Patent 12664970
SPEECH TRANSLATION WITH PERFORMANCE CHARACTERISTICS
3y 2m to grant Granted Jun 23, 2026
Patent 12626715
ROLE SEPARATION METHOD, ELECTRONIC DEVICE, AND COMPUTER STORAGE MEDIUM
3y 4m to grant Granted May 12, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

3-4
Expected OA Rounds
62%
Grant Probability
99%
With Interview (+37.5%)
2y 9m (~0m remaining)
Median Time to Grant
High
PTA Risk
Based on 32 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month