Prosecution Insights
Last updated: August 17, 2026
Application No. 18/745,098

Gesture Playback System

Final Rejection §101§103
Filed
Jun 17, 2024
Examiner
MASTERS, KRISTEN MICHELLE
Art Unit
2659
Tech Center
2600 — Communications
Assignee
Comcast Cable Communications LLC
OA Round
2 (Final)
65%
Grant Probability
Favorable
3-4
OA Rounds
10m
Est. Remaining
86%
With Interview

Examiner Intelligence

Grants 65% — above average
65%
Career Allowance Rate
32 granted / 49 resolved
+3.3% vs TC avg
Strong +21% interview lift
Without
With
+21.2%
Interview Lift
resolved cases with interview
Typical timeline
3y 0m
Avg Prosecution
22 currently pending
Career history
84
Total Applications
across all art units

Statute-Specific Performance

§101
37.9%
-2.1% vs TC avg
§103
48.1%
+8.1% vs TC avg
§102
7.8%
-32.2% vs TC avg
§112
3.4%
-36.6% vs TC avg
Black line = Tech Center average estimate • Based on career data from 49 resolved cases

Office Action

§101 §103
Detailed Action This communication is in response to the Arguments and Amendments filed on 4/13/2026. Claims 1-5 and 8-20 are pending and have been examined. Claims 6 and 7 have been cancelled. Claims 1-5 and 8-20 have been rejected. Hence this action has been made Final. Independent Claims 1, 12, and 17 are method claims, respectively. Apparent priority: 6/17/2024. Any previous objection/rejection not mentioned in this Office Action has been withdrawn by the Examiner. Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Response to Amendment The Applicant has amended the claims to include “1. (Currently Amended) A method comprising: accessing, by one or more computing devices, audio information associated with content; translating each phrase of a plurality of phrases, from the audio information, into a corresponding sign language gesture in a sequence of sign language gestures; determining an allocated duration for each gesture in the sequence of sign language gestures; determining, based on the allocated duration for each gesture, a plurality of gesture playback rate; and rates by, for each sign language gesture in the sequence of sign language gestures: determining a calculated gesture playback rate; comparing the calculated gesture playback rate and a corresponding minimum playback threshold; and selecting, as a gesture playback rate of the plurality of gesture playback rates, a larger of the calculated gesture playback rate and the corresponding minimum playback threshold; and providing, for output, data comprising generating, based on the plurality of gesture playback rates, a visual rendering, for display via a content player device, of the sequence of sign language gestures at the determined gesture playback rate.” Regarding the Rejections under 35 U.S.C. § 103 Applicant notes The cited references at least fail to disclose or suggest "comparing the calculated gesture playback rate and a corresponding minimum playback threshold" and "selecting, as a gesture playback rate of the plurality of gesture playback rates, a larger of the calculated gesture playback rate and the corresponding minimum playback threshold" of claim 1. Examiner notes Jawahar teaches [0042] In some embodiments, the system 100 includes a language model to control the language, an accent, and a duration of speech, and includes an alignment module that aligns the input speech with the pose sequence, thereby eliminating false pose sequences…. [0065] FIG. 5 is an exemplary representation 500 of a sign language video that is automatically generated for an input speech according to some embodiments herein. The exemplary representation 500 includes different pose sequences at 502 and one or more spectrograms at 504. The pose sequences at 502 are matched with corresponding one or more spectrograms at 504 to generate the sign language video using the machine learning model. (examiner notes the playback rate is included in the alignment of the spectrogram to the pose sequence) and dynamic time warping is used to calculate and select the best alignment [0064] In some embodiments, quality of the generated sign language pose sequences can be evaluated using Dynamic Time Warping (DTW) and Probability of Correct Key points (PCK) scores. The DTW may find an optimal alignment between two time series by non-linearly warping the pose sequences. The PCK may be used in pose detections and generation to evaluate the probability of pose key points to be close to the ground truth key points. Applicant notes Jawahar generally describes "a method for automatically generating a sign language video from an input speech using a machine learning model" (Jawahar, paragraph [0006]). With reference to the "minimum playback threshold" as recited in dependent claim 6, the Office Action cites paragraph [0066] of Jawahar, which states, "generating one or more pose sequences for a current time step of the plurality of spectrograms using a first machine learning model.. . automatically generating, using a second machine learning model, a sign language video for the input speech using the one or more pose sequences and the plurality of spectrograms when the one or more pose sequences are matched with corresponding the plurality of spectrograms that are extracted." Examiner notes Jawahar relates to teaching or communicating with deaf persons, animation and gesture generation. Jawahar describes [0002] Embodiments of this disclosure generally relate to automatically generating a sign language video using a machine learning model, and more particularly, to a system and method for automatically generating a sign language video from an input speech using the machine learning model with improved sign gestures and speaker's emotion. Applicant notes However, Jawahar fails to disclose or suggest "comparing the calculated gesture playback rate and a corresponding minimum playback threshold" or "selecting, as a gesture playback rate of the plurality of gesture playback rates, a larger of the calculated gesture playback rate and the corresponding minimum playback threshold" of claim 1. Notably, Jawahar is silent about a minimum playback threshold. Examiner notes See Jawahar Fig. 5 and [0064] In some embodiments, quality of the generated sign language pose sequences can be evaluated using Dynamic Time Warping (DTW) and Probability of Correct Key points (PCK) scores. The DTW may find an optimal alignment between two time series by non-linearly warping the pose sequences. The PCK may be used in pose detections and generation to evaluate the probability of pose key points to be close to the ground truth key points Applicant notes Additionally, Natesan fails to cure the deficient disclosure of Jawahar at least because Natesan fails to disclose or suggest a "minimum playback threshold," much less "selecting, as a gesture playback rate of the plurality of gesture playback rates, a larger of the calculated gesture playback rate and the corresponding minimum playback threshold" of claim 1. Examiner notes Natesan relates to sign language translation, transforming into visible information, and phrasal analysis. Natesan teaches [0014] …. Further, for each sentence, along with the sentence, the sentence start and end times may be identified in the speech video. The sentence start and end times may be utilized to ensure that the sign language video that is generated and the original speech video play in sync based, for example, on alignment of each sentence start and end time. Examiner has considered the arguments in regards to the prior art of Natesan and finds it persuasive. Hence, Natesan has been withdrawn. Regarding the Rejections under 35 U.S.C. §101 Applicant notes The Office characterizes the claim limitations as mental processes that can be performed by a human. Specifically, the Office alleges that the previously recited features of claim 1 "relate[ ] to a human using hand movements to output sign language gestures at a rate" (Office Action, pp. 2-3). Applicant respectfully submits that the Office's characterization is improper. The August 4, 2025 USPTO Memorandum titled "Reminders on evaluating subject matter eligibility of claims under 35 U.S.C. 101" ("Memorandum") admonishes that "[t]he mental process grouping is not without limits" and that "[e]xaminers are reminded not to expand this grouping in a manner that encompasses claim limitations that cannot practically be performed in the human mind" (Memorandum, p. 2). Specifically, the Memorandum establishes that "a claim does not recite a mental process when it contains limitation(s) that cannot practically be performed in the human mind, for instance when the human mind is not equipped to perform the claim limitation(s)" (Memorandum, p. 2) (emphases added). Applicant submits that claim 1 recites several limitations that cannot practically be performed in the human mind, including but not limited to "generating, based on the plurality of gesture playback rates, a visual rendering, for display via a content player device, of the sequence of sign language gestures." Notably, a human mind alone is not equipped to generate a visual rendering of a sequence of sign language gestures for display via a content player device. This limitation is analogous to the MPEP's example of "a claim to a method for rendering a halftone image of a digital image by comparing, pixel by pixel, the digital image against a blue noise mask, where the method required the manipulation of computer data structures (e.g., the pixels of a digital image and a two-dimensional array known as a mask) and the output of a modified computer data structure (a halftoned digital image), Research Corp. Techs., 627 F.3d at 868, 97 USPQ2d at 1280" (MPEP § 2106.04(a)(2)) (emphasis in original). Just as rendering a halftone image cannot practically be performed in the human mind, generating a visual rendering of a sequence of sign language gestures for display via a content player device cannot practically be performed in the human mind. Accordingly, claim 1 does not recite a mental process. Examiner notes The August 4, 2025 USPTO Memo’s reminder that the mental process grouping has limits is acknowledged, but it does not preclude application of the mental process grouping here because several claim limitations recite high level cognitive and physical manipulation concepts and processes that a human can perform (e.g., accessing audio information, translating phrases, gesturing, determining the speed at which gesturing is to be performed) that can be characterized as mental/data processing concepts. The fact that there is a display component does not automatically remove them from the judicial exception analysis without claim language or specification evidence showing a concrete technological improvement to computer functionality. Similarly, claim 12 recites "adjusting, based on the plurality of gesture playback rates, a playback speed of a visual rendering, displayed via a content player device, of the sequence of sign language gestures," which cannot practically be performed in the human mind because a human cannot adjust a playback speed of a visual rendering displayed via a content player device. In addition, Examiner notes a display device is noted as an additional limitation. A human can adjust a speed at which translation into sign language occurs. claim 17 recites "generating a visual rendering, for display via a content player device, of the sequence of sign language gestures by: applying the first gesture playback rate to the first sign language gesture in the visual rendering; and applying the second gesture playback rate to the second sign language gesture in the visual rendering," which similarly cannot practically be performed in the human mind because a human cannot generate a visual rendering by applying gesture playback rates to sign language gestures in the visual rendering. Accordingly, claims 12 and 17 also do not recite mental processes. Examiner notes a display device is noted as an additional limitation. A human can adjust a speed at which translation into sign language occurs. Even assuming, arguendo, that the claims recite an abstract idea, the claims integrate any such alleged abstract idea into a practical application under Step 2A Prong Two. MPEP § 2106.04(d) identifies that claims integrate a judicial exception into a practical application when they include "[a]n improvement in the functioning of a computer, or an improvement to other technology or technical field." The MPEP further explains that "the presence of a non-physical or intangible additional element does not doom the claims, because tangibility is not necessary for eligibility under the Alice/Mayo test," citing McRO, Inc. v. Bandai Namco Games Am. Inc., 837 F.3d 1299, 1315, 120 USPQ2d 1091, 1102 (Fed. Cir. 2016), which held "that a process producing an intangible result (a sequence of synchronized, animated characters) was eligible because it improved an existing technological process" (MPEP § 2106.04(d)). Present claims 1, 12, and 17 similarly produce synchronized sign language gesture renderings-an improvement to accessibility technology that synchronizes sign language gesture output with audio content. Examiner notes applicant points to (McRO, Inc. v. Bandai Namco Games Am) and cases recognizing technical improvements from objective rules to improve technological processes. Examiner notes “generating a visual rendering, the sequence of sign language gestures” is not the same as “produce synchronized sign language gesture renderings” Examiner notes applicant would need to amend the claims, tying the claims to specific rules, showing exact mathematical, logical or conditional ruleset running the software, and clearly articulate the improvement, showing practical application of the invention. As written, the claim’s high level functional language does not make the improvement evident on the face of the claim; therefore the rejection stands absent amendment. Examiner has considered the arguments with regards to the Natesan prior art and finds it persuasive. Hence, Natesan has been withdrawn. Applicant's arguments and amendments are persuasive and overcome the under 35 U.S.C. 103 rejection. Hence, new grounds of rejection have been made in view of Jawahar (U.S. Patent Number US 20230290371 A1) in view of TAYEBI (U.S. Patent Number US 20240203017 A1). Claim Rejections - 35 USC § 101 35 U.S.C. 101 reads as follows: Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title. Claims 1-5 and 8-20 are rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more. Independent claim 1 recites, “1. (Currently Amended) A method comprising: accessing, by one or more computing devices, audio information associated with content; [this relates to a human using language and reasoning in the mind to assess audio information with content] translating each phrase of a plurality of phrases, from the audio information, into a corresponding sign language gesture in a sequence of sign language gestures; (This relates to a human using natural language translation and to translate audio into sign language.) determining a plurality of gesture playback rates by, for each sign language gesture in the sequence of sign language gestures: (This relates to a human using perception to determine a duration for each sign language gesture.) determining a calculated gesture playback rate; (This relates to a human using perception to determine a rate.) comparing the calculated gesture playback rate and a corresponding minimum playback threshold; and [this relates to a human comparing playback rates against a threshold in the human mind.] selecting, as a gesture playback rate of the plurality of gesture playback rates, a larger of the calculated gesture playback rate and the corresponding minimum playback threshold; and [this relates a to a human selecting a larger gesture playback rate.] generating, based on the plurality of gesture playback rates, a visual rendering, for display via a content player device, of the sequence of sign language gestures.” (This relates to a human generating a rendering using pen and paper.) a device is noted as an additional limitation. The Dependent Claims do not include additional limitations that could incorporate the abstract idea into a practical application or cause the Claim as a whole to amount to significantly more than the underlying abstract idea. Regarding Independent Claim 12, claim 12 is a method claim with limitations similar to that of Claim 1 and is rejected under the same rational. no additional elements. Regarding Independent Claim 17, claim 17 is a method claim with limitations similar to that of Claim 1 and is rejected under the same rational. no additional elements. Dependent claim 2 recites, “2. The method of claim 1, further comprising storing a text segment associated with the audio information, a start time associated with the text segment, and a duration associated with the text segment. (This relates to a human using pen and paper to store a text segment and start time and a duration.) no additional elements. Dependent claim 3 recites, “3. The method of claim 1, further comprising providing, for output, data comprising: a start time associated with an allocated duration of each sign language gesture in the sequence of sign language gestures; and an end time associated with the allocated duration of each gesture. (This relates to a human using perception to note a start time and duration.) no additional elements. Dependent claim 4 recites, “4. The method of claim 1, further comprising: sending a text segment associated with the audio information; accessing a start time associated with the text segment and a segment duration associated with the text segment; (This relates to a human using pen and paper to store a text segment and start time and a duration.) and determining, for each sign language gesture in the sequence of sign language gestures: an allocated duration by dividing the segment duration with a total number of gestures in the sequence of sign language gestures.” (This relates to a human counting gestures.) no additional elements. Dependent claim 5 recites, “5. The method of claim 1, wherein the determining the calculated gesture playback rate comprises, for each sign language gesture in the sequence of sign language gestures, determining the calculated gesture playback rate by dividing a predetermined gesture time with the allocated duration. (This relates to a human using perception to determine a rate and using logic and reasoning dividing gesture time with duration.) no additional elements. Dependent claim 8 recites, “8. The method of claim 1, further comprising: receiving a content player rate; determining that the content player rate is above a normal rate; (This relates to human using perception to determine a rate is above normal.) and determining, for each sign language gesture in the sequence of sign language gestures, an adjusted gesture playback rate by multiplying the content player rate with the calculated gesture playback rate. (this relates to a human applying logic and reasoning to determine an adjusted rate using mathematical calculations.) no additional elements. Dependent claim 9 recites, “9. The method of claim 1, further comprises the sequence of sign language gestures are associated with American Sign Language. (This relates to a human performing sign language using natural language understanding and gestures) no additional elements. Dependent claim 10 recites, “10. The method of claim 1, further comprises: determining, based on a context of a text segment, an intensity associated with each gesture. (This relates to a human determining intensity of gestures using perception and natural language understanding) no additional elements. Dependent claim 11 recites, “11. The method of claim 1, wherein the translating further comprises training a machine learning model to translate the audio information to the sequence of sign language gestures. (This relates to a human using natural language understanding to translate audio into gestures.) no additional elements. As to claim 13, claim 13 is a parallel method claim with limitations similar to that of claim 6 and is rejected under the same rationale. Dependent claim 14 recites, 14. (Currently Amended) The method of claim 13, further comprising, for each sign language gesture in the sequence of sign language gestures: determining, the calculated gesture playback rate by dividing a predetermined gesture time with the allocated duration. (This relates to a human performing a mathematical calculation.) no additional elements. As to claim 15, claim 15 is a parallel method claim with limitations similar to that of claim 2and is rejected under the same rationale. As to claim 16, claim 16 is a parallel method claim with limitations similar to that of claim 3 and is rejected under the same rationale. Dependent claim 18 recites, “18. (Currently Amended) The method of claim 17, further comprising: determining, the first calculated gesture playback rate by dividing a first predetermined gesture time, associated with the first sign language gesture, with a first allocated duration associated with the first sign language gesture; (This relates to a human determining a rate according to a threshold using perception and logic and reasoning) and determining the second calculated gesture playback rate by dividing a second predetermined gesture time, associated with the second sign language gesture, with a second allocated duration associated with the second sign language gesture. (This relates to a mathematical calculation a human can perform.) no additional elements. As to claim 19, claim 19 is a parallel method claim with limitations similar to that of claim 2 and is rejected under the same rationale. As to claim 20, claim 20 is a parallel method claim with limitations similar to that of claim 3 and is rejected under the same rationale. Claim Rejections - 35 USC § 103 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claims 1-5, 8-12, 17, 19 and 20 are rejected under 35 U.S.C. 103 as being unpatentable over Jawahar (U.S. Patent Number US 20230290371 A1) in view of TAYEBI (U.S. Patent Number US 20240203017 A1). Regarding independent Claim 1, Jawahar teaches 1. A method comprising: accessing, by one or more computing devices, audio information associated with content; translating each phrase of a plurality of phrases, from the audio information into a corresponding sign language gesture in a sequence of sign language gestures; (see Jawahar [0006] “In view of the foregoing, an embodiment herein provides a method for automatically generating a sign language video from an input speech using a machine learning model. The method includes extracting a plurality of spectrograms of an input speech by (i) encoding, using an encoder, a time domain series of the input speech to a frequency domain series, and (ii) decoding, using a decoder, a plurality of tokens for time steps of the frequency domain series. Each spectrogram comprises at least one visual representation of a strength of the input speech over time, the input speech is obtained from a user device associated with a user. The method includes generating a plurality of pose sequences for a current time step of the plurality of spectrograms using a first machine learning model. The first machine learning model is trained by correlating historical pose sequences of historical users in historical sign language videos with historical spectrograms of historical input speeches. The method includes automatically generating, using a discriminator of a second machine learning model, a sign language video for the input speech using the plurality of pose sequences and the plurality of spectrograms when the plurality of pose sequences are matched with corresponding the plurality of spectrograms that are extracted.”) determining, a plurality of gesture playback rates by; (see Jawahar [0066] “FIG. 6 illustrates a flow diagram of a method for automatically generating a sign language video from an input speech using the machine learning model of FIG. 1 according to some embodiments herein. At the step 602, the method includes extracting a plurality of spectrograms of an input speech by (i) encoding, using an encoder, a time domain series of the input speech to a frequency domain series, and (ii) decoding, using a decoder, a plurality of tokens for time steps of the frequency domain series. Each spectrogram comprises at least one visual representation of a strength of the input speech over time, wherein the input speech is obtained from a user device associated with a user. At the step 604, the method includes generating one or more pose sequences for a current time step of the plurality of spectrograms using a first machine learning model. The first machine learning model is trained by correlating historical pose sequences of historical users in historical sign language videos with historical spectrograms of historical input speeches. At the step 606, the automatically generating, using a second machine learning model, a sign language video for the input speech using the one or more pose sequences and the plurality of spectrograms when the one or more pose sequences are matched with corresponding the plurality of spectrograms that are extracted.”) generating, based on the plurality of gesture playback rates, a visual rendering, for display via a content player device, of the sequence of sign language gestures.(see Jawahar [0056] “The generator 308 is configured to input the Mel spectrogram of the input speech from the speech encoder. In some embodiments, the generator 308 embeds the Mel Spectrogram accurately. The generator 308 may include a positional encoder that encodes the Mel Spectrogram according to the previous time step and the current time step. The generator 308 is configured to input the predicted sign poses from the pose decoder. In some embodiments, the generator 308 embeds the predicted sign poses accurately. The generator 308 may include a positional encoder that encodes the predicted sign poses according to the previous time step and the current time step.”) (see Jawahar [0057] “The generator 308 associates with the speech embedding and the pose embedding, that is configured to learn attention aware representations for the modalities. In some embodiments, the modalities include any of the input speech and the pose sequence. The generator 308 is configured to fuse the modalities that learns to embed the speech segments of the input speech into the pose sequence. In some embodiments, the generator 308 learns a relationship between the input speech and the pose sequence. The generator 308 may merge the speech segments with the pose sequence. The generator 308 is configured to apply attention to the fused embedding of two modalities and to find whether the two modalities match or not.”) (see Jawahar [0071]…generated sign language video.”) Jawahar does not specifically teach for each sign language gesture in the sequence of sign language gestures: determining a calculated gesture playback rate; However, TAYEBI does teach this limitation (see TAYEBI [0186-0187] In some examples, the speed at which each animation is played may be scaled, and/or the speed at which all the animation play is scaled. [0187] In some examples, scaling may be undertaken before the transition animation is generated, as scaling the adjacent animations may change the kinematic boundary conditions. comparing the calculated gesture playback rate and a corresponding minimum playback threshold; (see TAYEBI [0188-0190] For example, a first animation time of the first animation A.sub.1, a second animation time of second animation A.sub.2 and transition time for the transition animation A.sub.t are scaled based on a global time variable.[0189] Additionally, or alternatively, a first animation time of the first animation A.sub.1 is scaled based on a first animation time scaling variable.[0190] Additionally, or alternatively, a second animation time of the second animation A.sub.2 is scaled based on a second animation time scaling variable. and selecting, as a gesture playback rate of the plurality of gesture playback rates, a larger of the calculated gesture playback rate and the corresponding minimum playback threshold; and (see TAYEBI [0191] “Additionally, or alternatively, a transition time of the transition animation A.sub.t is scaled based on a transition time animation scaling variable.”) [0080] In another aspect, there is provided a method for generating an animated sentence in sign language with said method consisting of the following steps: Using motion capture to create animations for a dictionary of words; cleaning and standardizing the animations; selecting the required animations or typing the required words to retrieve the required words; generating a transition animation between each word to seamlessly link the words together; generate a transition for the fingers based on the distance to next sign; determining the transition time of facial animations to synchronize with body animations; overlaying post-processing animation effects onto the character.”) Jawahar and TAYEBI are in the same field of endeavor of signal processing, therefore, it would have been prima facie obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the method of Jawahar to incorporate for each sign language gesture in the sequence of sign language gestures: determining a calculated gesture playback rate; comparing the calculated gesture playback rate and a corresponding minimum playback threshold;. and selecting, as a gesture playback rate of the plurality of gesture playback rates, a larger of the calculated gesture playback rate and the corresponding minimum playback threshold of TAYEBI. This allows improved realism of an animation by reproducing the life-like subtleties in the movement of motion-captured animation as recognized by TAYEBI [0126]. As to Independent Claim 12, Claim 12 is a parallel method claim with limitations similar to that of claim 1 and is rejected under the same rationale. As to Independent Claim 17, Claim 17 is a parallel method claim with limitations similar to that of claim 1 and is rejected under the same rationale. Furthermore TAYEBI teaches determining a first gesture playback rate for the first sign language gesture by selecting a larger of a first calculated gesture playback rate and a first minimum playback threshold; determining a second gesture playback rate for the second sign language gesture (see TAYEBI [0137] “In some examples, the first animation A.sub.1 is completed before transitioning to the second animation A.sub.2, however in other examples Animation A.sub.1 does not need to complete to initiate a transition. If the transition is initiated while an animation is playing or midway through a transition, a buffer, tracking the history of a character's position (for example bone locations and/or joint rotations may be used). This may be used in cases where the transitions are required to be dynamic and won't be guaranteed to play in an uninterrupted sequence.”) by selecting a larger of a second calculated gesture playback rate and a second minimum playback threshold; and generating a visual rendering, for display via a content player device, of the sequence of sign language gestures by: applying the first gesture playback rate to the first sign language gesture in the visual rendering; and applying the second gesture playback rate to the second sign language gesture in the visual rendering. (see TAYEBI [0188-0190] For example, a first animation time of the first animation A.sub.1, a second animation time of second animation A.sub.2 and transition time for the transition animation A.sub.t are scaled based on a global time variable.[0189] Additionally, or alternatively, a first animation time of the first animation A.sub.1 is scaled based on a first animation time scaling variable.[0190] Additionally, or alternatively, a second animation time of the second animation A.sub.2 is scaled based on a second animation time scaling variable.”)(see TAYEBI [0191] “Additionally, or alternatively, a transition time of the transition animation A.sub.t is scaled based on a transition time animation scaling variable.”) [0080] In another aspect, there is provided a method for generating an animated sentence in sign language with said method consisting of the following steps: Using motion capture to create animations for a dictionary of words; cleaning and standardizing the animations; selecting the required animations or typing the required words to retrieve the required words; generating a transition animation between each word to seamlessly link the words together; generate a transition for the fingers based on the distance to next sign; determining the transition time of facial animations to synchronize with body animations; overlaying post-processing animation effects onto the character.”) Jawahar and TAYEBI are in the same field of endeavor of signal processing, therefore, it would have been prima facie obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the method of Jawahar to incorporate determining a first gesture playback rate for the first sign language gesture by selecting a larger of a first calculated gesture playback rate and a first minimum playback threshold; determining a second gesture playback rate for the second sign language gesture by selecting a larger of a second calculated gesture playback rate and a second minimum playback threshold; and generating a visual rendering, for display via a content player device, of the sequence of sign language gestures by: applying the first gesture playback rate to the first sign language gesture in the visual rendering; and applying the second gesture playback rate to the second sign language gesture in the visual rendering of TAYEBI. This allows improved realism of an animation by reproducing the life-like subtleties in the movement of motion-captured animation as recognized by TAYEBI [0126]. As to Claim 2, Jawahar in view of TAYEBI teaches 2. The method of claim 1, Furthermore, Jawahar teaches further comprising storing a text segment associated with the audio information, a start time associated with the text segment, and a duration associated with the text segment. (see Jawahar [0054] “FIG. 3 is a block diagram of the second machine learning model 108B of FIG. 1 according to some embodiments herein. The second machine learning model 108B includes a ground truth module 302, a discriminator 304, a loss function 306, and a generator 308. The generator 308 generates the plurality of pose sequences, and the sign language video for the input speech. The ground truth module 302 determines ground truth spectrograms, ground truth pose sequences, ground truth input speech. The discriminator 304 discriminates between at least one of ground truth spectrograms, ground truth pose sequences, ground truth input speech, and generated plurality of pose sequences, generated sign language video. The loss function 306 is generated when the discriminator 304 discriminates between at least one of ground truth spectrograms, ground truth pose sequences, ground truth input speech, and generated plurality of pose sequences, generated sign language video.”) (see Jawahar [0055] “The discriminator 304 is configured to match the speech segments with the pose sequence. The pose sequence may be predicted sign poses. In some embodiments, the discriminator 304 includes a separate speech and pose embedding layers that learns a high dimensional embedding of the input speech and the pose sequence.”) (see Jawahar [0056] “The generator 308 is configured to input the Mel spectrogram of the input speech from the speech encoder. In some embodiments, the generator 308 embeds the Mel Spectrogram accurately. The generator 308 may include a positional encoder that encodes the Mel Spectrogram according to the previous time step and the current time step. The generator 308 is configured to input the predicted sign poses from the pose decoder. In some embodiments, the generator 308 embeds the predicted sign poses accurately. The generator 308 may include a positional encoder that encodes the predicted sign poses according to the previous time step and the current time step.”) As to Claim 3, Jawahar in view of TAYEBI teaches 3. The method of claim 1, Furthermore, Jawahar teaches further comprising: comprising providing, for output, data comprising: a start time associated with an allocated duration of each sign language gesture in the sequence of sign language gestures; and an end time associated with the allocated duration (see Jawahar [0042] “In some embodiments, the system 100 includes a language model to control the language, an accent, and a duration of speech, and includes an alignment module that aligns the input speech with the pose sequence, thereby eliminating false pose sequences. The database may store instructions to generate the sign language video with the input speech. The system 100 may include a memory that stores instructions and a processor that executes the stored instructions to generate the sign language video with the input speech.”) (see Jawahar [0063] “The system 100 may use an embedding size of d.sub.model=512, N=2 layers and number of heads, M=8. In some embodiments, the system 100 uses Xavier initialization and Adam optimizer with an initial learning rate of 10-3 for training the multi-task transformer and the cross modal discriminator. Data augmentations like predicting multiple-frame poses may be determined at each time step. In some embodiments, the system 100 predict 10 frames at every time step to penalize the network heavily for producing mean pose sequences.”) (see Jawahar [0064] “In some embodiments, quality of the generated sign language pose sequences can be evaluated using Dynamic Time Warping (DTW) and Probability of Correct Key points (PCK) scores. The DTW may find an optimal alignment between two time series by non-linearly warping the pose sequences. The PCK may be used in pose detections and generation to evaluate the probability of pose key points to be close to the ground truth key points..”) (see Jawahar [0012] “The method includes generating a plurality of pose sequences for a current time step of the plurality of spectrograms using a first machine learning model. The first machine learning model is trained by correlating historical pose sequences of historical users in historical sign language videos with historical spectrograms of historical input speeches.”) As to Claim 4, Jawahar in view of TAYEBI teaches 4. The method of claim 1, Furthermore, Jawahar further comprising: sending a text segment associated with the audio information; accessing a start time associated with the text segment and a segment duration associated with the text segment; and determining, for each sign language gesture in the sequence of sign language gestures, an allocated duration by dividing the segment duration with a total number of gestures in the sequence of sign language gestures. (see Jawahar [0066] “FIG. 6 illustrates a flow diagram of a method for automatically generating a sign language video from an input speech using the machine learning model of FIG. 1 according to some embodiments herein. At the step 602, the method includes extracting a plurality of spectrograms of an input speech by (i) encoding, using an encoder, a time domain series of the input speech to a frequency domain series, and (ii) decoding, using a decoder, a plurality of tokens for time steps of the frequency domain series. Each spectrogram comprises at least one visual representation of a strength of the input speech over time, wherein the input speech is obtained from a user device associated with a user. At the step 604, the method includes generating one or more pose sequences for a current time step of the plurality of spectrograms using a first machine learning model. The first machine learning model is trained by correlating historical pose sequences of historical users in historical sign language videos with historical spectrograms of historical input speeches. At the step 606, the automatically generating, using a second machine learning model, a sign language video for the input speech using the one or more pose sequences and the plurality of spectrograms when the one or more pose sequences are matched with corresponding the plurality of spectrograms that are extracted.”) As to Claim 5, Jawahar in view of TAYEBI teaches 5. The method of claim 1, Furthermore, Jawahar teaches wherein the determining the calculated gesture playback rate comprises, for each sign language gesture in the sequence of sign language gestures: determining, calculated gesture playback rate by dividing a predetermined gesture time with the allocated duration. (see Jawahar [0054] “FIG. 3 is a block diagram of the second machine learning model 108B of FIG. 1 according to some embodiments herein. The second machine learning model 108B includes a ground truth module 302, a discriminator 304, a loss function 306, and a generator 308. The generator 308 generates the plurality of pose sequences, and the sign language video for the input speech. The ground truth module 302 determines ground truth spectrograms, ground truth pose sequences, ground truth input speech. The discriminator 304 discriminates between at least one of ground truth spectrograms, ground truth pose sequences, ground truth input speech, and generated plurality of pose sequences, generated sign language video. The loss function 306 is generated when the discriminator 304 discriminates between at least one of ground truth spectrograms, ground truth pose sequences, ground truth input speech, and generated plurality of pose sequences, generated sign language video.”) (see Jawahar [0055] The discriminator 304 is configured to match the speech segments with the pose sequence. The pose sequence may be predicted sign poses. In some embodiments, the discriminator 304 includes a separate speech and pose embedding layers that learns a high dimensional embedding of the input speech and the pose sequence.”) (see Jawahar [0056] “The generator 308 is configured to input the Mel spectrogram of the input speech from the speech encoder. In some embodiments, the generator 308 embeds the Mel Spectrogram accurately. The generator 308 may include a positional encoder that encodes the Mel Spectrogram according to the previous time step and the current time step. The generator 308 is configured to input the predicted sign poses from the pose decoder. In some embodiments, the generator 308 embeds the predicted sign poses accurately. The generator 308 may include a positional encoder that encodes the predicted sign poses according to the previous time step and the current time step.”) As to Claim 8, Jawahar in view of TAYEBI teaches 8. The method of claim 1, Furthermore, Jawahar teaches, further comprising: receiving a content player rate; determining that the content player rate is above a normal rate; and determining, for each sign language gesture in the sequence of sign language gestures, an adjusted gesture playback rate by multiplying the content player rate with the calculated gesture playback rate. (see Jawahar [0063] “The system 100 may use an embedding size of d.sub.model=512, N=2 layers and number of heads, M=8. In some embodiments, the system 100 uses Xavier initialization and Adam optimizer with an initial learning rate of 10-3 for training the multi-task transformer and the cross modal discriminator. Data augmentations like predicting multiple-frame poses may be determined at each time step. In some embodiments, the system 100 predict 10 frames at every time step to penalize the network heavily for producing mean pose sequences.”) (see Jawahar [0064] “In some embodiments, quality of the generated sign language pose sequences can be evaluated using Dynamic Time Warping (DTW) and Probability of Correct Key points (PCK) scores. The DTW may find an optimal alignment between two time series by non-linearly warping the pose sequences. The PCK may be used in pose detections and generation to evaluate the probability of pose key points to be close to the ground truth key points.”) (see Jawahar [0066] “FIG. 6 illustrates a flow diagram of a method for automatically generating a sign language video from an input speech using the machine learning model of FIG. 1 according to some embodiments herein. At the step 602, the method includes extracting a plurality of spectrograms of an input speech by (i) encoding, using an encoder, a time domain series of the input speech to a frequency domain series, and (ii) decoding, using a decoder, a plurality of tokens for time steps of the frequency domain series. Each spectrogram comprises at least one visual representation of a strength of the input speech over time, wherein the input speech is obtained from a user device associated with a user. At the step 604, the method includes generating one or more pose sequences for a current time step of the plurality of spectrograms using a first machine learning model. The first machine learning model is trained by correlating historical pose sequences of historical users in historical sign language videos with historical spectrograms of historical input speeches. At the step 606, the automatically generating, using a second machine learning model, a sign language video for the input speech using the one or more pose sequences and the plurality of spectrograms when the one or more pose sequences are matched with corresponding the plurality of spectrograms that are extracted.”) As to Claim 9, Jawahar in view of TAYEBI teaches 9. The method of claim 1, Furthermore, Jawahar teaches, wherein the sequence of sign language gestures are associated with American Sign Language. (see Jawahar [0066] “FIG. 6 illustrates a flow diagram of a method for automatically generating a sign language video from an input speech using the machine learning model of FIG. 1 according to some embodiments herein. At the step 602, the method includes extracting a plurality of spectrograms of an input speech by (i) encoding, using an encoder, a time domain series of the input speech to a frequency domain series, and (ii) decoding, using a decoder, a plurality of tokens for time steps of the frequency domain series. Each spectrogram comprises at least one visual representation of a strength of the input speech over time, wherein the input speech is obtained from a user device associated with a user. At the step 604, the method includes generating one or more pose sequences for a current time step of the plurality of spectrograms using a first machine learning model. The first machine learning model is trained by correlating historical pose sequences of historical users in historical sign language videos with historical spectrograms of historical input speeches. At the step 606, the automatically generating, using a second machine learning model, a sign language video for the input speech using the one or more pose sequences and the plurality of spectrograms when the one or more pose sequences are matched with corresponding the plurality of spectrograms that are extracted.”)(examiner notes see Figure 5 for American Sign language) As to Claim 10, Jawahar in view of TAYEBI teaches 10. The method of claim 1, Furthermore, TAYEBI teaches further comprising: determining, based on a context of a text segment associated with the audio information, an intensity associated with each sign language gesture in the sequence of sign language gestures. (see TAYEBI [0130] “Once the animations match the intended movements of the source signer, a standardization process is followed. The spine within the humanoid skeleton is standardized across all animations with a neutral animation as during motion capture, the signer's spine may shift quite significantly between words. This unnecessary movement needs to be removed while preserving any natural spine movement during a sign or intended movement that is part of the sign. Next, the velocity of an animation can be normalized using a mapping function that takes extreme ends of the velocity of a sign and maps it to an average value/range. This prevents any fluctuations in speed between recorded signs causing large changes in sign speed between two sequential words. These animations should then be stored in a database with tags specifying attributes about the sign such as but not limited to sign one-handedness, sign sentiment, and sign intensity.”) Jawahar and TAYEBI are in the same field of endeavor of signal processing, therefore, it would have been prima facie obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the method of Jawahar and TAYEBI to incorporate further comprising: determining, based on a context of a text segment associated with the audio information, an intensity associated with each sign language gesture in the sequence of sign language gestures. of TAYEBI. This allows improved realism of an animation by reproducing the life-like subtleties in the movement of motion-captured animation as recognized by TAYEBI [0126]. As to Claim 11, Jawahar in view of TAYEBI teaches 11. The method of claim 1, Furthermore, Jawahar teaches wherein the translating comprises: training a machine learning model to translate the audio information to the sequence of sign language gestures. (see Jawahar [0006] "In view of the foregoing, an embodiment herein provides a method for automatically generating a sign language video from an input speech using a machine learning model. The method includes extracting a plurality of spectrograms of an input speech by (i) encoding, using an encoder, a time domain series of the input speech to a frequency domain series, and (ii) decoding, using a decoder, a plurality of tokens for time steps of the frequency domain series. Each spectrogram comprises at least one visual representation of a strength of the input speech over time, the input speech is obtained from a user device associated with a user. The method includes generating a plurality of pose sequences for a current time step of the plurality of spectrograms using a first machine learning model. The first machine learning model is trained by correlating historical pose sequences of historical users in historical sign language videos with historical spectrograms of historical input speeches. The method includes automatically generating, using a discriminator of a second machine learning model, a sign language video for the input speech using the plurality of pose sequences and the plurality of spectrograms when the plurality of pose sequences are matched with corresponding the plurality of spectrograms that are extracted.”) As to claim 19, claim 19 is a parallel method claim with limitations similar to that of claim 2 and is rejected under the same rationale. As to claim 20, claim 20 is a parallel method claim with limitations similar to that of claim 3 and is rejected under the same rationale. Claims 13-16 and 18 are rejected under 35 U.S.C. 103 as being unpatentable over Jawahar (U.S. Patent Number US 20230290371 A1) in view of TAYEBI (U.S. Patent Number US 20240203017 A1) and further in view of in view of WANG (U.S. Patent Number US 20230326369 A1). As to claim 13 Jawahar in view of TAYEBI teach 13. The method of claim 12, Jawahar in view of TAYEBI do not specifically teach wherein the determining the corresponding gesture playback rate, for each sign language gesture in the sequence of sign language gestures, comprises: determining that a calculated gesture playback rate is less than the corresponding minimum playback threshold; and determining the corresponding gesture playback rate to be equivalent to the corresponding minimum playback threshold However, Wang does teach this limitation (see Wang [0191] “In some embodiments, the listener text contain corresponding timestamps, and the timestamps are used for indicating a time interval of audio corresponding to the listener text on an audio time axis. The extraction module 1302 is configured to determine candidate clip durations of candidate sign language video clips corresponding to the candidate compression statements; determine audio clip durations of audio corresponding to the text statements based on timestamps corresponding to the text statements; and determine, based on the candidate clip durations and the audio clip durations, the target compression statements from the candidate compression statements through the dynamic path planning algorithm, a video time axis of sign language video corresponding to texts composed of the target compression statements being aligned with the audio time axis of the audio corresponding to the listener text.”) (see Wang [0197] “In some embodiments, the acquisition module 1301 is configured to: acquire the input listener text; acquire a subtitle file, and extract the listener text from the subtitle file; acquire an audio file, perform speech recognition on the audio file to obtain a speech recognition result, and generate the listener text based on the speech recognition result; and acquire a video file, perform character recognition on video frames of the video file to obtain a character recognition result, and generate the listener text based on the character recognition result.”) (see Wang [0198] In summary, in the embodiments of this application, the summary text are obtained by performing text summarization extraction on the listener text, and then the text length of the listener text are shortened, so that the finally generated sign language video can keep synchronization with the audio corresponding to the listener text. Since the sign language video is generated based on the sign language text after the summary text are converted into the sign language text conforming to the grammatical structures of a hearing-impaired person, the sign language video can better express the content to a hearing-impaired person, improving the accuracy of the sign language video.”) Jawahar in view of TAYEBI and Wang are in the same field of endeavor of signal processing, therefore, it would have been prima facie obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the method of combination of Jawahar and TAYEBI to incorporate receiving a minimum playback threshold, wherein the slow gesture playback rate is less than the minimum playback threshold; and adjusting, the slow gesture playback rate to be equivalent to the minimum playback threshold of Wang. This allows for improved generation efficiency of the sign language video as recognized by Wang [0013-0014]. As to claim 14 Jawahar in view of TAYEBI teach 14. The method of claim 12, Jawahar in view of TAYEBI do not specifically teach further comprising, for each sign language gesture in the sequence of sign language gestures: determining, the calculated gesture playback rate by dividing a predetermined gesture time with the allocated duration However, Wang does teach this limitation (see Wang [0191] “In some embodiments, the listener text contain corresponding timestamps, and the timestamps are used for indicating a time interval of audio corresponding to the listener text on an audio time axis. The extraction module 1302 is configured to determine candidate clip durations of candidate sign language video clips corresponding to the candidate compression statements; determine audio clip durations of audio corresponding to the text statements based on timestamps corresponding to the text statements; and determine, based on the candidate clip durations and the audio clip durations, the target compression statements from the candidate compression statements through the dynamic path planning algorithm, a video time axis of sign language video corresponding to texts composed of the target compression statements being aligned with the audio time axis of the audio corresponding to the listener text.”) (see Wang [0197] “In some embodiments, the acquisition module 1301 is configured to: acquire the input listener text; acquire a subtitle file, and extract the listener text from the subtitle file; acquire an audio file, perform speech recognition on the audio file to obtain a speech recognition result, and generate the listener text based on the speech recognition result; and acquire a video file, perform character recognition on video frames of the video file to obtain a character recognition result, and generate the listener text based on the character recognition result.”) (see Wang [0198] In summary, in the embodiments of this application, the summary text are obtained by performing text summarization extraction on the listener text, and then the text length of the listener text are shortened, so that the finally generated sign language video can keep synchronization with the audio corresponding to the listener text. Since the sign language video is generated based on the sign language text after the summary text are converted into the sign language text conforming to the grammatical structures of a hearing-impaired person, the sign language video can better express the content to a hearing-impaired person, improving the accuracy of the sign language video.”) Jawahar in view of TAYEBI and Wang are in the same field of endeavor of signal processing, therefore, it would have been prima facie obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the method of combination of Jawahar and TAYEBI to incorporate determining, for each gesture, the slow gesture playback rate by dividing a predetermined gesture time with the allocated duration, wherein a minimum playback threshold is greater than the slow gesture playback rate; and adjusting, based on the determination, the slow gesture playback rate to be equivalent to the minimum playback threshold of Wang. This allows for improved generation efficiency of the sign language video as recognized by Wang [0013-0014]. As to claim 15 Jawahar in view of TAYEBI teach 15. The method of claim 12, Jawahar in view of TAYEBI do not specifically teach further comprising sending a text segment associated with the dialogue, a start time associated with the text segment, and a duration associated with the text segment However, Wang does teach this limitation (see Wang [0191] “In some embodiments, the listener text contain corresponding timestamps, and the timestamps are used for indicating a time interval of audio corresponding to the listener text on an audio time axis. The extraction module 1302 is configured to determine candidate clip durations of candidate sign language video clips corresponding to the candidate compression statements; determine audio clip durations of audio corresponding to the text statements based on timestamps corresponding to the text statements; and determine, based on the candidate clip durations and the audio clip durations, the target compression statements from the candidate compression statements through the dynamic path planning algorithm, a video time axis of sign language video corresponding to texts composed of the target compression statements being aligned with the audio time axis of the audio corresponding to the listener text.”) (see Wang [0197] “In some embodiments, the acquisition module 1301 is configured to: acquire the input listener text; acquire a subtitle file, and extract the listener text from the subtitle file; acquire an audio file, perform speech recognition on the audio file to obtain a speech recognition result, and generate the listener text based on the speech recognition result; and acquire a video file, perform character recognition on video frames of the video file to obtain a character recognition result, and generate the listener text based on the character recognition result.”) (see Wang [0198] In summary, in the embodiments of this application, the summary text are obtained by performing text summarization extraction on the listener text, and then the text length of the listener text are shortened, so that the finally generated sign language video can keep synchronization with the audio corresponding to the listener text. Since the sign language video is generated based on the sign language text after the summary text are converted into the sign language text conforming to the grammatical structures of a hearing-impaired person, the sign language video can better express the content to a hearing-impaired person, improving the accuracy of the sign language video.”) Jawahar in view of TAYEBI and Wang are in the same field of endeavor of signal processing, therefore, it would have been prima facie obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the method of combination of Jawahar and TAYEBI to incorporate sending a text segment associated with the audio information, a start time associated with the text segment, and a duration associated with the text segment of Wang. This allows for improved generation efficiency of the sign language video as recognized by Wang [0013-0014]. As to claim 16 Jawahar in view of TAYEBI teach 16. The method of claim 12, Jawahar in view of TAYEBI do not specifically teach further comprising providing, for output, data comprising: a start time associated with the allocated duration of each sign language gesture in the sequence of sign language gestures; and an end time associated with the allocated duration. However, Wang does teach this limitation (see Wang [0191] “In some embodiments, the listener text contain corresponding timestamps, and the timestamps are used for indicating a time interval of audio corresponding to the listener text on an audio time axis. The extraction module 1302 is configured to determine candidate clip durations of candidate sign language video clips corresponding to the candidate compression statements; determine audio clip durations of audio corresponding to the text statements based on timestamps corresponding to the text statements; and determine, based on the candidate clip durations and the audio clip durations, the target compression statements from the candidate compression statements through the dynamic path planning algorithm, a video time axis of sign language video corresponding to texts composed of the target compression statements being aligned with the audio time axis of the audio corresponding to the listener text.”) (see Wang [0197] “In some embodiments, the acquisition module 1301 is configured to: acquire the input listener text; acquire a subtitle file, and extract the listener text from the subtitle file; acquire an audio file, perform speech recognition on the audio file to obtain a speech recognition result, and generate the listener text based on the speech recognition result; and acquire a video file, perform character recognition on video frames of the video file to obtain a character recognition result, and generate the listener text based on the character recognition result.”) (see Wang [0198] In summary, in the embodiments of this application, the summary text are obtained by performing text summarization extraction on the listener text, and then the text length of the listener text are shortened, so that the finally generated sign language video can keep synchronization with the audio corresponding to the listener text. Since the sign language video is generated based on the sign language text after the summary text are converted into the sign language text conforming to the grammatical structures of a hearing-impaired person, the sign language video can better express the content to a hearing-impaired person, improving the accuracy of the sign language video.”) Jawahar in view of TAYEBI and Wang are in the same field of endeavor of signal processing, therefore, it would have been prima facie obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the method of combination of Jawahar and TAYEBI to incorporate wherein the data further comprises: a start time associated with the allocated duration of each gesture; and an end time associated with the allocated duration of each gesture of Wang. This allows for improved generation efficiency of the sign language video as recognized by Wang [0013-0014]. As to claim 18 Jawahar in view of TAYEBI teach 18. The method of claim 17, Jawahar in view of TAYEBI do not specifically teach further comprising: determining, the first calculated gesture playback rate by dividing a first predetermined gesture time, associated with the first sign language gesture, with the a first allocated duration associated with the first sign language gesture; and determining the second calculated gesture playback rate by dividing a second predetermined gesture time, associated with the second sign language gesture, with a second allocated duration associated with the second sign language gesture. However, Wang does teach this limitation (see Wang [0191] “In some embodiments, the listener text contain corresponding timestamps, and the timestamps are used for indicating a time interval of audio corresponding to the listener text on an audio time axis. The extraction module 1302 is configured to determine candidate clip durations of candidate sign language video clips corresponding to the candidate compression statements; determine audio clip durations of audio corresponding to the text statements based on timestamps corresponding to the text statements; and determine, based on the candidate clip durations and the audio clip durations, the target compression statements from the candidate compression statements through the dynamic path planning algorithm, a video time axis of sign language video corresponding to texts composed of the target compression statements being aligned with the audio time axis of the audio corresponding to the listener text.”) (see Wang [0197] “In some embodiments, the acquisition module 1301 is configured to: acquire the input listener text; acquire a subtitle file, and extract the listener text from the subtitle file; acquire an audio file, perform speech recognition on the audio file to obtain a speech recognition result, and generate the listener text based on the speech recognition result; and acquire a video file, perform character recognition on video frames of the video file to obtain a character recognition result, and generate the listener text based on the character recognition result.”) (see Wang [0198] In summary, in the embodiments of this application, the summary text are obtained by performing text summarization extraction on the listener text, and then the text length of the listener text are shortened, so that the finally generated sign language video can keep synchronization with the audio corresponding to the listener text. Since the sign language video is generated based on the sign language text after the summary text are converted into the sign language text conforming to the grammatical structures of a hearing-impaired person, the sign language video can better express the content to a hearing-impaired person, improving the accuracy of the sign language video.”) Jawahar in view of TAYEBI and Wang are in the same field of endeavor of signal processing, therefore, it would have been prima facie obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the method of combination of Jawahar and TAYEBI to incorporate determining, for each gesture, the gesture playback rate by dividing a predetermined gesture time with the allocated duration; receiving the content player rate, wherein the content play rate is above a normal rate; and determining, an adjusted gesture playback rate by multiplying the content player rate with the gesture playback rate. of Wang. This allows for improved generation efficiency of the sign language video as recognized by Wang [0013-0014]. Conclusion Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a). A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action. Any inquiry concerning this communication or earlier communications from the examiner should be directed to KRISTEN MICHELLE MASTERS whose telephone number is (703)756-1274. The examiner can normally be reached M-F 8:30 AM - 5:00 PM. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Pierre Louis Desir can be reached at 571-272-7799. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /KRISTEN MICHELLE MASTERS/Examiner, Art Unit 2659 /PIERRE LOUIS DESIR/Supervisory Patent Examiner, Art Unit 2659
Read full office action

Prosecution Timeline

Jun 17, 2024
Application Filed
Jan 13, 2026
Non-Final Rejection mailed — §101, §103
Apr 13, 2026
Response Filed
Jul 02, 2026
Final Rejection mailed — §101, §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12707198
ACOUSTIC ECHO CANCELLATION SYSTEM AND ASSOCIATED METHOD
3y 0m to grant Granted Aug 11, 2026
Patent 12694889
PROFANITY FILTER FOR COLLABORATION SESSIONS IN HETEROGENOUS COMPUTING PLATFORMS
3y 10m to grant Granted Jul 28, 2026
Patent 12664366
CROSS-DOMAIN LABEL-ADAPTIVE STANCE DETECTION
3y 11m to grant Granted Jun 23, 2026
Patent 12592219
Hearing Device User Communicating With a Wireless Communication Device
4y 5m to grant Granted Mar 31, 2026
Patent 12548569
METHOD AND SYSTEM OF DETECTING AND IMPROVING REAL-TIME MISPRONUNCIATION OF WORDS
3y 2m to grant Granted Feb 10, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

3-4
Expected OA Rounds
65%
Grant Probability
86%
With Interview (+21.2%)
3y 0m (~10m remaining)
Median Time to Grant
Moderate
PTA Risk
Based on 49 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month