Prosecution Insights
Last updated: October 02, 2026
Application No. 18/892,737

SIGN LANGUAGE IN GAMES

Final Rejection §101§103
Filed
Sep 23, 2024
Priority
Sep 28, 2023 — GB 2314903.2
Examiner
GILLS, KURTIS
Art Unit
Tech Center
Assignee
Sony Group Corporation
OA Round
2 (Final)
58%
Grant Probability
Moderate
3-4
OA Rounds
1y 6m
Est. Remaining
87%
With Interview

Examiner Intelligence

Grants 58% of resolved cases
58%
Career Allowance Rate
327 granted / 565 resolved
-2.1% vs TC avg
Strong +29% interview lift
Without
With
+29.2%
Interview Lift
resolved cases with interview
Typical timeline
3y 7m
Avg Prosecution
29 currently pending
Career history
600
Total Applications
across all art units

Statute-Specific Performance

§101
38.5%
-1.5% vs TC avg
§103
43.8%
+3.8% vs TC avg
§102
6.6%
-33.4% vs TC avg
§112
6.7%
-33.3% vs TC avg
Black line = Tech Center average estimate • Based on career data from 565 resolved cases

Office Action

§101 §103
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . DETAILED ACTION Notice to Applicant In response to the communication received on 07/27/2026, the following is a Final Office Action for Application No. 18892737. Status of Claims Claims 20, 23, 25-29, 31-38, and 40-41 are pending. Claims 1-19, 21-22, 24, 30, and 39 are cancelled. Priority As required by M.P.E.P. 201.14(c), acknowledgement is made of applicant’s claim for priority based on: 18892737 filed 09/23/2024, claims foreign priority to 2314903.2, filed 09/28/2023. Response to Amendments Applicant’s amendments have been fully considered. Applicant’s amendments to the claims overcome the 35 U.S.C 112 rejection and hence the 35 U.S.C. 112 rejection has been withdrawn. Applicant’s amendments to the claims overcome the 35 U.S.C 101 rejection with respect to non-transitory issues and hence the 35 U.S.C. 101 rejection with respect to non-transitory issues has been withdrawn. Response to Arguments Applicant’s arguments with respect to the claims have been considered but are moot in light of the new grounds of rejection, as necessitated by amendment. As per the 101 rejection, Applicant argues that the claims are in favor of eligibility per Prong One of Step 2A, however Examiner respectfully disagrees. Per Prong One of Step 2A, the identified recitation of an abstract idea falls within at least one of the Abstract Idea Groupings consisting of: Mathematical Concepts, Mental Processes, or Certain Methods of Organizing Human Activity. Particularly, the identified recitation falls within the Mental Processes including concepts performed in the human mind (including an observation, evaluation judgment, opinion) and/or Certain Methods of Organizing Human Activity including managing personal behavior or relationships or interactions between people (including social activities, teaching, and following rules of instructions). Since the recitation of the claims falls into at least one of the above Groupings, there is a basis for providing further analysis with regard to Prong Two of Step 2A to determine whether the recitation of an abstract idea is deduced to being directed to an abstract idea. Thus, the rejection is maintained. Applicant argues that the claims are in favor of eligibility per Prong Two of Step 2A, however Examiner respectfully disagrees. Per Prong Two of Step 2A, this judicial exception is not integrated into a practical application because the claim as a whole does not integrate the identified abstract idea into a practical application. The computer, processor and/or memory medium is recited at a high level of generality, i.e., as a generic processor performing a generic computer function of processing/transmitting data. This generic processor server limitation is no more than mere instructions to apply the exception using a generic computer component. Further, computer, processor and/or memory medium to inter alia perform the function of outputting the image for display is mere instruction to apply an exception using a generic computer component which cannot integrate a judicial exception into a practical application. Accordingly, this/these additional element(s) does/do not integrate the abstract idea into a practical application because it does not impose any meaningful limits on practicing the abstract idea. In other words, the present claims use a generic processing device and memory medium to inter alia perform the function of outputting the image for display which is a concept that can be performed in the human mind. The processor is merely used to perform the function(s), and the processor does not integrate the abstract idea into a practical application since there are no meaningful limits on practicing the abstract idea. Thus, since the claims are directed to the determined judicial exception in view of the two prongs of Step 2A, the 2019 PEG flowchart is directed to Step 2B. Thus, the rejection is maintained. Applicant argues that the claims are in favor of eligibility per Step 2B, however Examiner respectfully disagrees. Therein, the additional elements and combinations therewith are examined in the claims to determine whether the claims as a whole amounts to significantly more than the judicial exception. It is noted here that the additional elements are to be considered both individually and as an ordered combination. In this case, the claims each at most comprise additional elements of: computer, processor and/or memory medium. Taken individually, the additional limitations each are generically recited and thus does not add significantly more to the respective limitations. Further, computer, processor and/or memory medium to inter alia perform the function of outputting the image for display is mere instruction to apply an exception using a generic computer component which cannot provide an inventive concept in Step 2B (or, looking back to Step 2A, cannot integrate a judicial exception into a practical application). For further support, the Applicant’s specification supports the claims being directed to use of a generic computer/memory type structure. Taken as an ordered combination, the claim(s) does/do not include additional elements that are sufficient to amount to significantly more than the judicial exception because the limitations are directed to limitations referenced in Alice Corp. that are not enough to qualify as significantly more when recited in a claim with an abstract idea include the non-limiting or non-exclusive examples of MPEP § 2106.05. Thus, the rejection is maintained. In an effort to further expedite prosecution, see: July 2024 Subject Matter Eligibility Examples, Example 47. Anomaly Detection. Per the analysis of claim 2 Example 47, the analysis refers to MPEP 2106.05(f) which provides the following considerations for determining whether a claim simply recites a judicial exception with the words “apply it” (or an equivalent), such as mere instructions to implement an abstract idea on a computer: (1) whether the claim recites only the idea of a solution or outcome i.e., the claim fails to recite details of how a solution to a problem is accomplished; (2) whether the claim invokes computers or other machinery merely as a tool to perform an existing process; and (3) the particularity or generality of the application of the judicial exception. Although the additional elements, e.g. (per Example 47) “using a trained ANN”, limits the identified judicial exceptions, e.g. (per Example 47) “detecting one or more anomalies in a data set using the trained ANN” and, e.g. (per Example 47) “analyzing the one or more detected anomalies using the trained ANN to generate anomaly data,” this type of limitation merely confines the use of the abstract idea to a particular technological environment, e.g. (per Example 47: neural networks) and thus fails to add an inventive concept to the claims. See MPEP 2106.05(h). As an exemplary direction for claim limitations to be eligible, see claims 1 and 3 of Example 47. Claim Rejections - 35 USC § 101 35 U.S.C. 101 reads as follows: Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title. Claims 20, 23, 25-29, 31-38, and 40-41 are rejected under 35 U.S.C. 101 because the claimed invention is directed to a judicial exception (i.e., a law of nature, a natural phenomenon, or an abstract idea) without significantly more. The claims fall within statutory class of process or machine or manufacture; hence, the claims fall under statutory category of Step 1. Step 2 is the two-part analysis from Alice Corp. (also called the Mayo test). The 2019 PEG makes two changes in Step 2A: It sets forth new procedure for Step 2A (called “revised Step 2A”) under which a claim is not “directed to” a judicial exception unless the claim satisfies a two-prong inquiry. The two-prong inquiry is as follows: Prong One: evaluate whether the claim recites a judicial exception (an abstract idea enumerated in the 2019 PEG, a law of nature, or a natural phenomenon). If claim recites an exception, then Prong Two: evaluate whether the claim recites additional elements that integrate the exception into a practical application of the exception. The claim(s) recite(s) the following abstract idea indicated by non-boldface font and additional limitations indicated by boldface font: 20. A computer-implemented method comprising: generating, based at least on a plurality of game assets, an image of a scene of a virtual environment; determining, based at least on the plurality of game assets, that the scene includes a character that is communicating using sign language and that a hand of the character is not visible in the image; in response to determining that the scene includes a character that is communicating using sign language and that the hand of the character is not visible in the image, adjusting generation of the image , wherein adjusting generation of the image comprises at least one of: adjusting a virtual viewpoint of the scene to cause the hand of the character to be visible in the generated image, or moving the character from a first location in the virtual environment to a second location in the virtual environment at which the hand of the character is visible within the generated image; and outputting the image for display. [or] 37. A system comprising: one or more processors, anode or more non-transitory computer-readable media that store instructions which, when executed by the one or more processors, cause the one or more processors to perform operations comprising: generating, based at least on a plurality of game assets, an image of a scene of a virtual environment; determining, based at least on the plurality of game assets, that the scene includes a character that is communicating using sign language and that a hand of the character is not visible in the image; in response to determining that the scene includes a character that is communicating using sign language and that the hand of the character is not visible in the image, adjusting generation of the image , wherein adjusting generation of the image comprises at least one of: adjusting a virtual viewpoint of the scene to cause the hand of the character to be visible in the generated image, or moving the character from a first location in the virtual environment to a second location in the virtual environment at which the hand of the character is visible within the generated image; and outputting the image for display. [or] 41. One or more non-transitory computer-readable media that store instructions which, when executed by one or more processors, cause the one or more processors to perform operations comprising: generating, based at least on a plurality of game assets, an image of a scene of a virtual environment; determining, based at least on the plurality of game assets, that the scene includes a character that is communicating using sign language and that a hand of the character is not visible in the image; in response to determining that the scene includes a character that is communicating using sign language and that the hand of the character is not visible in the image, adjusting generation of the image , wherein adjusting generation of the image comprises at least one of: adjusting a virtual viewpoint of the scene to cause the hand of the character to be visible in the generated image, or moving the character from a first location in the virtual environment to a second location in the virtual environment at which the hand of the character is visible within the generated image; and outputting the image for display. The claim(s) recite(s) the following summarization of the abstract idea which includes processing a game state of a game and to generate, based on a plurality of game assets, an image of a scene executed by the additional element(s) of computer readable storage medium, processing circuitry, output circuitry and/or processor. This falls into at least the Abstract Idea Grouping of Mental Processes since the information can be analyzed by an abstract evaluation judgment process. Thus, per Prong One of Step 2A, the identified recitation of an abstract idea falls within at least one of the Abstract Idea Groupings consisting of: Mathematical Concepts, Mental Processes, or Certain Methods of Organizing Human Activity since the identified recitation falls within the Mental Processes including concepts performed in the human mind (including an observation, evaluation judgment, opinion). Per Prong Two of Step 2A, this judicial exception is not integrated into a practical application because the claim as a whole does not integrate the identified abstract idea into a practical application. The processing circuitry, output circuitry, processors, and/or computer-readable media is recited at a high level of generality, i.e., as a generic processor performing a generic computer function of processing/transmitting data. This generic processors, and/or computer-readable media limitation is no more than mere instructions to apply the exception using a generic computer component. Further, outputting the image for display by a processors, and/or computer-readable media is mere instruction to apply an exception using a generic computer component which cannot integrate a judicial exception into a practical application. Accordingly, this/these additional element(s) does/do not integrate the abstract idea into a practical application because it does not impose any meaningful limits on practicing the abstract idea. The machine-learned model in claims 33-36 is used to generally apply the abstract idea without placing any limits on how the machine-learned model functions. Rather, these limitations only recite the outcome of the functions and do not include any details about how the functions via the machine-learned model are accomplished. See MPEP 2106.05(f). Thus, since the claims are directed to the determined judicial exception in view of the two prongs of Step 2A, the 2019 PEG flowchart is directed to Step 2B. Per Step 2B, the additional elements and combinations therewith are examined in the claims to determine whether the claims as a whole amounts to significantly more than the judicial exception. It is noted here that the additional elements are to be considered both individually and as an ordered combination. In this case, the claims each at most comprise additional elements of: computer, processors, and computer-readable media. Taken individually, the additional limitations each are generically recited and thus does not add significantly more to the respective limitations. The additional element of a machine-learned model is at best mere instructions to “apply” the abstract ideas, which cannot provide an inventive concept. See MPEP 2106.05(f). Further, outputting the image for display by a processing circuitry, output circuitry, processors, and/or computer-readable media is mere instruction to apply an exception using a generic computer component which cannot provide an inventive concept in Step 2B (or, looking back to Step 2A, cannot integrate a judicial exception into a practical application). For further support, the Applicant’s specification supports the claims being directed to use of a generic computer/memory type structure at ¶0033 wherein “the method could be performed by a general purpose computer under the control of a computer program comprising instructions which, when executed by a computer, cause the computer to perform the above method.” Taken as an ordered combination, the claim(s) does/do not include additional elements that are sufficient to amount to significantly more than the judicial exception because the limitations are directed to limitations referenced in Alice Corp. that are not enough to qualify as significantly more when recited in a claim with an abstract idea include, as a non-limiting or non-exclusive examples: i. Adding the words "apply it" (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, e.g., a limitation indicating that a particular function such as creating and maintaining electronic records is performed by a computer, as discussed in Alice Corp., 134 S. Ct. at 2360, 110 USPQ2d at 1984 (see MPEP § 2106.05(f)); PNG media_image1.png 18 19 media_image1.png Greyscale ii. Simply appending well-understood, routine, conventional activities previously known to the industry, specified at a high level of generality, to the judicial exception, e.g., a claim to an abstract idea requiring no more than a generic computer to perform generic computer functions that are well-understood, routine and conventional activities previously known to the industry, as discussed in Alice Corp., 134 S. Ct. at 2359-60, 110 USPQ2d at 1984 (see MPEP § 2106.05(d)); PNG media_image1.png 18 19 media_image1.png Greyscale iii. Adding insignificant extra-solution activity to the judicial exception, e.g., mere data gathering in conjunction with a law of nature or abstract idea such as a step of obtaining information about credit card transactions so that the information can be analyzed by an abstract mental process, as discussed in CyberSource v. Retail Decisions, Inc., 654 F.3d 1366, 1375, 99 USPQ2d 1690, 1694 (Fed. Cir. 2011) (see MPEP § 2106.05(g)); or PNG media_image1.png 18 19 media_image1.png Greyscale v. Generally linking the use of the judicial exception to a particular technological environment or field of use, e.g., a claim describing how the abstract idea of hedging could be used in the commodities and energy markets, as discussed in Bilski v. Kappos, 561 U.S. 593, 595, 95 USPQ2d 1001, 1010 (2010) or a claim limiting the use of a mathematical formula to the petrochemical and oil-refining fields, as discussed in Parker v. Flook. The courts have recognized the following computer functions inter alia to be well-understood, routine, and conventional functions when they are claimed in a merely generic manner: performing repetitive calculations; receiving, processing, and storing data (e.g., the present claims); electronically scanning or extracting data; electronic recordkeeping; automating mental tasks (e.g., process/machine/manufacture for performing the present claims); and receiving or transmitting data (e.g., the present claims). The dependent claims do not cure the above stated deficiencies, and in particular, the dependent claims further narrow the abstract idea without reciting additional elements other than the addressed machine learned model of claims 33-36 that integrate the exception into a practical application of the exception or providing significantly more than the abstract idea. Since there are no elements or ordered combination of elements that amount to significantly more than the judicial exception, the claims are not eligible subject matter under 35 USC §101. Thus, viewed as a whole, these additional claim element(s) do not provide meaningful limitation(s) to transform the abstract idea into a patent eligible application of the abstract idea such that the claim(s) amounts to significantly more than the abstract idea itself. Therefore, the claim(s) are rejected under 35 U.S.C. 101 as being directed to non-statutory subject matter. Claim Rejections - 35 USC § 103 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102 of this title, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows: 1. Determining the scope and contents of the prior art. 2. Ascertaining the differences between the prior art and the claims at issue. 3. Resolving the level of ordinary skill in the pertinent art. 4. Considering objective evidence present in the application indicating obviousness or nonobviousness. Claims 20, 23, 25-29, 31-38, and 40-41 are rejected under 35 U.S.C. 103 as being unpatentable over Marey et al. (US 20220343576 A1) hereinafter referred to as Marey in view of Adamo-Villani et al. (US 20060134585 A1) hereinafter referred to as Adamo-Villani. Marey teaches: Claim 20. A computer-implemented method comprising: generating, based at least on a plurality of game assets, an image of a scene of a virtual environment (¶0038 As referred to herein, processing circuitry should be understood to mean circuitry based on one or more microprocessors, microcontrollers, digital signal processors, programmable logic devices, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), etc., and may include a multi-core processor (e.g., dual-core, quad-core, hexa-core, or any suitable number of cores ¶0048 Avatar generation module 234 may generate a virtual avatar performing signs translated by text-to-sign language module 232. Avatar generation module 234 may use a Virtual Reality (VR) rendering technique, such as the Unreal® Engine, to generate a virtual avatar that mimics the user or character's movement and expression. Unreal® Engine is a software development environment suite for building virtual and augmented reality graphics, game development, architectural visualization, content creation, broadcast, or any other real-time applications. Avatar generation module 234 includes an expression reconstruction module 238 and a motion reconstruction module 240.); determining, based at least on the plurality of game assets, that the scene includes a character that is communicating using sign language and that a hand of the character is not visible in the image (¶0069 The translation application performs sentiment analysis for each character displayed in the content item using emotion-detection techniques described above. In contrast to FIG. 3, exemplary user interface 400 of FIG. 4 shows a different emotional state (“happy”) of the character. In a romantic scene of the movie, when Sally is speaking the line “I love you,” the translation application infers that Sally is probably happy based on the context of the movie and the words spoken by Sally. Therefore, the happy feeling is exhibited on the face of Sally's avatar 402 (e.g., smiling with teeth and shiny eyes). ¶0074 FIG. 6 depicts an exemplary embodiment 600 for processing signs in real time for interacting with a non-sign user, in accordance with some embodiments of the disclosure. In exemplary embodiment 600, a sign user 602 interacts with a clerk 606 at a hair salon using sign user's device 604. If sign user 602 speaks sign language (“I would like to cancel my appointment”), then a clerk's device 608 captures a gesture of sign user 602 speaking the sign language via a camera of clerk's device 608. Based on the captured gesture (e.g., left hand pointing to right wrist), a translation application running on clerk's device 608 translates the sign language and converts the sign language into text. The translated text is displayed on clerk's device 608 or is output as audio via a speaker of clerk's device 608.); in response to determining that the scene includes a character that is communicating using sign language and that the hand of the character is not visible in the image, adjusting generation of the image , wherein adjusting generation of the image comprises at least one of:adjusting a virtual viewpoint of the scene to cause the hand of the character to be visible in the generated image, or moving the character from a first location in the virtual environment to a second location in the virtual environment at which the hand of the character is visible within the generated image (¶0070 In some embodiments, the appearance of the avatars resembles the appearance of the characters in the content item. As shown in exemplary interface 400, female avatar 402 is displayed as wearing the same clothes as the female character, such as the same hairband, dress, or necklace. This allows the user to feel like the actual character in the content item is speaking the sign language, resulting in a more active engagement with the content item. ¶0009 The translation application animates the movement of hands, fingers, and facial expressions of an avatar by changing the relative positions of the hands, fingers, arms, or parts of the face of the avatar. For example, an avatar has one or more skeleton models that refer to different parts of the body. A hand model includes references to the different fingers (e.g., index finger, middle finger, ring finger). An arm model references different parts of the arm (e.g., above the elbow, below the elbow). A face model includes references to different parts of the face, such as the left eye, right eye, forehead, lips, etc. Based on the images or videos showing gestures of corresponding signs stored in the database, the transformation of the avatar is performed. The transformation (e.g., up, down, pitch roll, etc.) is applied to corresponding parts of the body using one or more skeleton models. For example, based on the images or videos that include a movement of certain parts of the body, the translation application identifies the moving parts of the body and identifies one or more relevant skeleton models that are required to perform the sign ¶0081 The transformation of the avatar based on the user input is achieved by identifying relevant skeleton models underlying the avatar structure to perform the command. As shown in FIG. 8, the “buy” sign involves moving the arms and the fingers. One or more skeleton models are applied to the avatar for performing the “buy” sign, and each skeleton model includes references to the joints. As shown by an avatar displayed in exemplary user interface 800, a hand skeleton model includes references to the different wrists 804 (e.g., left wrist, right wrist). An arm model references different joints of the arm 802 (e.g., above the elbow, below the elbow); and outputting the image for display (¶0074 FIG. 6 depicts an exemplary embodiment 600 for processing signs in real time for interacting with a non-sign user, in accordance with some embodiments of the disclosure. In exemplary embodiment 600, a sign user 602 interacts with a clerk 606 at a hair salon using sign user's device 604. If sign user 602 speaks sign language (“I would like to cancel my appointment”), then a clerk's device 608 captures a gesture of sign user 602 speaking the sign language via a camera of clerk's device 608. Based on the captured gesture (e.g., left hand pointing to right wrist), a translation application running on clerk's device 608 translates the sign language and converts the sign language into text. The translated text is displayed on clerk's device 608 or is output as audio via a speaker of clerk's device 608. ¶0084 The user may request a presentation of a content item (e.g., “romantic movie”) via a translation application running on computing device 114. In response to the request, at step 902, control circuitry 202 receives a content item from a content item source 106. The content item contains a video component and an audio component. The audio component includes a first plurality of spoken words (e.g., “I love you”) in a first language (e.g., English). The video component includes a character (e.g., Sally in FIG. 3) who speaks the first plurality of spoken words in the first language.). Although not explicitly taught by Marey, Adamo-Villani teaches in the analogous art of interactive animation system for sign language: a hand of the character is not visible in the image (¶0046 We note that different views of the signer's hands and arms while signing are necessary for effective practice and learning. For example, the front and two side views are the views generally observed in conversation, while the point of view of the signer is useful when learning how to sign. In order to acquire proficiency in signing it is important to be able to observe one's own hands and arms in the process of producing the correct signs. For this reason we have provided a tumble tool which allows a 360 degrees rotation of the camera around the signer. FIG. 6 shows the camera controls pop-up window with two different views of the signer. Another point worth noting is the ability to control the speed of signing. For the beginner (i.e., a hearing parent of a deaf child) observing people signing at natural speed, the signs usually cannot be resolved--the motion of the fingers appearing as a moving blur. {the 360 degree rotation includes a view whereby a hand of the character is not visible in the image}); adjusting a virtual viewpoint of the scene to cause the hand of the character to be visible in the generated image, or moving the character from a first location in the virtual environment to a second location in the virtual environment at which the hand of the character is visible within the generated image (¶0044 Each screen presents a consistent layout and visual style. In some examples, the screen layout may include two frames, as shown in FIG. 4. In the example shown in FIG. 4, the frame on the left is used to select the grade (k-1, 2 or 3) and/or the type of activity. The frame on the right is occupied by the 3D signer (FIG. 4). The upper area of the frame on the left (in green) is used to give textual feedback on the current activity, the bottom area contains the navigational buttons. The frame on the right contains a white text box, right below the signer, used to show the answer (in mathematical symbols) to the current problem. Below the answer box there is a camera icon and a slider represented by an arrow. The slider is used to control the speed of signing, the camera button opens a popup menu used to zoom in/out on the 3D signer, change the point of view and pan to the left or to the right within the 3D signer window. In other examples, the screen layout may include three frames, as shown in FIG. 5. With this layout, the tasks of learning and testing are more clearly separated. More importantly, the avatar is now placed in the middle of the screen and activity areas, such as the learning and testing activities, are placed on the left and right sides, respectively. The viewer can now fully attend to the center of the screen and use her peripheral vision to see the other two frames (containing the buttons/activities). Instead of a constant shift of gaze between the avatar and the buttons/activities, the spatial relationship between the activity areas is such that the user can focus on one of the activity areas while still understanding the signed communication from the avatar.). It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to combine the interactive animation system for sign language of Adamo-Villani with the sentiment-based interactive avatar system for sign language of Marey for the following reasons: (1) a finding that there was some teaching, suggestion, or motivation, either in the references themselves or in the knowledge generally available to one of ordinary skill in the art, to modify the reference or to combine reference teachings, e.g. Marey ¶0003 teaches that it is desirable for a deaf person to be able to focus on reading text and not miss a visual component, such as emotions that are expressed by a character in the audio-visual content (e.g., facial expression).; (2) a finding that there was reasonable expectation of success since the only difference between the claimed invention and the prior art being the lack of actual combination of the elements in a single prior art reference, e.g. Marey Abstract teaches a system f or doing presenting an avatar that speaks sign language based on sentiment of a speaker, and Adamo-Villani Abstract teaches system for interactive communication in sign language using computer animation; and (3) whatever additional findings based on the Graham factual inquiries may be necessary, in view of the facts of the case under consideration, to explain a conclusion of obviousness, e.g. Marey at least the above cited paragraphs, and Adamo-Villani at least the inclusively cited paragraphs. Therefore, it would be obvious to one skilled in the art at the time of the invention to combine the interactive animation system for sign language of Adamo-Villani with the sentiment-based interactive avatar system for sign language of Marey. The rationale to support a conclusion that the claim would have been obvious is that "a person of ordinary skill in the art would have been motivated to combine the prior art to achieve the claimed invention and whether there would have been a reasonable expectation of success in doing so." DyStar Textilfarben GmbH & Co. Deutschland KG v. C.H. Patrick Co., 464 F.3d 1356, 1360, 80 USPQ2d 1641, 1645 (Fed. Cir. 2006). See MPEP 2143(G). Marey teaches: Claim 23. The method of claims 20, comprising:determining that the at least one character is not understandable without use of audible, clues,wherein the determining whether the scene includes at least one character that is communicating using sign language is in response to determining that the at least one character is not understandable without use of audible clues (¶0014 The present disclosure addresses one or more comprehension issues the deaf person may experience by providing graphical representations of a real-time live avatar that speaks sign language and exhibits an emotion of a speaker on a display screen of the computing device. The present disclosure adds significant solutions to the existing problems, such as having to perform a phonemic task and not being able to fully grasp the visual component of the audio-visual content. Thereby, the present disclosure allows the deaf person to consume the content item asset within a reasonable time, understand the emotions of the speaker, and facilitates direct communication with non-sign users, resulting in an improved communication or content item environment for the deaf person.). Marey teaches: Claim 25. The method of claim 20, wherein determining whether the scene includes at least one character that is communicating using sign language is based at least on at least one hand of the at least one character or alternative visual cues representing the communication of the at least one character (¶0003 The audio component includes one or more words spoken in a first language (e.g., English). The translation application translates the words of the first language (e.g., “I am happy”) into a first sign of the first sign language (e.g., American Sign Language). The translation application determines an emotional state of a character who speaks the words in the first language in the content item. For example, the emotional state may be determined by at least spoken words of the first language, vocal tone, facial expression, or body expression of the character in the video. The translation application generates an avatar that performs the first sign of the first sign language (e.g., performing a “happy” sign), exhibiting the previously determined emotional state (e.g., happy expression on the face of the avatar). The content item and the avatar are concurrently presented to the user for display.). Marey teaches: Claim 26. The method of claim 25, wherein, adjusting the generation of the image further comprises to including the alternative visual cues representing the communication of the character (¶0055 As referred to herein, the term “content item” should be understood to mean an electronically consumable user asset, such as an electronic version of a printed book, electronic television programming, as well as pay-per-view program, on-demand program (as in video-on-demand (VOD) system), Internet content (e.g., streaming content, downloadable content, Webcasts, etc.), video clip, audio, content information, picture, rotating image, document, playlist, website, article, book, article, newspaper, blog, advertisement, chat session, social content item, application, games, and/or any other content item or multi content item and/or combination of the same ¶0071 To minimize the amount of obstruction caused by the avatar, the translation application may determine a non-focus area of the video. A non-focus area is a portion of a frame of the displayed content where an avatar can be generated. The non-focus area features less important content, such as the background of the scene (e.g., forest or ocean). The translation application retrieves metadata of the displayed content and identifies a candidate non-focus area for each frame. For example, if a portion of frames include action scenes where the objects in the frame are rapidly moving, then a non-focus area may be changed accordingly. On the other hand, if a portion of frames includes static scenes as shown in exemplary user interface 300, an avatar is placed at a default location of the frame and may remain in the same location for a portion of the content item. An avatar may be generated in various display modes. As shown in exemplary user interface 300, an avatar may be displayed in Picture-In-Picture (PIP) display. An avatar may be in a multi-window mode that is separate from a window that contains the video, as shown in exemplary user interface 400. The window of the avatar may be pinned to any corner of the screen during the display). Marey teaches: Claim 27. The method of claim 26, wherein adjusting the generation of the image comprises including, as the alternative visual cues, subtitles representing the communication of the at least one character (¶0057 The translation application may convert the speech to text using speech-to-text module 230. If closed caption data is received with the content item, the speech-to-text conversion step may be skipped. If the closed caption data is not available, then the speech-to-text module first converts the anchor's speech to text using any speech recognition techniques or voice recognition techniques. In one example, the translation application uses the lipreading techniques to interpret the movements of the lips, face, or tongue of the speaker to decipher the speech. Because the content item is broadcast in real time (e.g., live news), the conversion may be performed in real time.). Marey teaches: Claim 28. The method of claim 26, wherein adjusting the generation of the image comprises including, as the alternative visual cues, an interpreting character performing the sign language (¶0058 Once the translation application converts the speech to text, the translation application uses text-to-sign language module 232 to translate the text to corresponding signs by querying sign language source 108. For example, the translation application queries sign language source 108 for each word in the text and identifies a corresponding sign. Sign language source 108 includes a sign language dictionary that contains several videos or images of sign language signs, fingerspelled words, or other common signs used within a particular country. Based on the corresponding gestures or motions contained in the videos or images, the translation application identifies a corresponding sign for the text.). Marey teaches: Claim 29. The method of claim 26, comprising:translating the sign language into a language other than the sign language,wherein the alternative visual cues comprises a translation of the sign language into the language other than the sign language (¶0047 Translation application server 104 includes a speech-to-text module 230, a text-to-sign language module 232, avatar generation module 234, or sentiment analysis module 236. Speech-to-text module 230 converts a speech received via microphone 226 of computing device 114 to text. Speech-to-text module 230 may implement any machine learning speech recognition or voice recognition techniques, such as Google® DeepMind, to decipher the speech of a user or a character in a content item. Text-to-sign language module 232 receives converted text generated by speech-to-text module 230 and translates the text to sign language (e.g., signs). The text can be translated to any sign language, such as American Sign Language, British Sign Language, or Spanish Sign Language. Text-to-sign language module 232 may utilize the sign language source 108 when translating to sign language.). Marey teaches: Claim 31. The method of claim 20, wherein adjusting the generation if the image is based on determining whether a sign language enhancement mode or a sign language non-enhancement mode is activated (¶0070 In some embodiments, the appearance of the avatars resembles the appearance of the characters in the content item. As shown in exemplary interface 400, female avatar 402 is displayed as wearing the same clothes as the female character, such as the same hairband, dress, or necklace. This allows the user to feel like the actual character in the content item is speaking the sign language, resulting in a more active engagement with the content item. To minimize the amount of obstruction caused by the avatar, the translation application may determine a non-focus area of the video. A non-focus area is a portion of a frame of the displayed content where an avatar can be generated. The non-focus area features less important content, such as the background of the scene (e.g., forest or ocean). The translation application retrieves metadata of the displayed content and identifies a candidate non-focus area for each frame.). Marey teaches: Claim 32. The method of claim 2, wherein:when the sign language enhancement mode, determining that the scene comprises the character communicating using sign language is based at least on determining whether communication by the character as displayed in the image to be understandable without use of audible cues; andwhen the sign language non-enhancement mode, determining that the scene comprises the at least one character communicating using sign language not based at least on determining whether communication by the at least one character as displayed in the image to be understandable without use of audible cues (¶0071 To minimize the amount of obstruction caused by the avatar, the translation application may determine a non-focus area of the video. A non-focus area is a portion of a frame of the displayed content where an avatar can be generated. The non-focus area features less important content, such as the background of the scene (e.g., forest or ocean). The translation application retrieves metadata of the displayed content and identifies a candidate non-focus area for each frame. For example, if a portion of frames include action scenes where the objects in the frame are rapidly moving, then a non-focus area may be changed accordingly. On the other hand, if a portion of frames includes static scenes as shown in exemplary user interface 300, an avatar is placed at a default location of the frame and may remain in the same location for a portion of the content item. An avatar may be generated in various display modes. As shown in exemplary user interface 300, an avatar may be displayed in Picture-In-Picture (PIP) display. An avatar may be in a multi-window mode that is separate from a window that contains the video, as shown in exemplary user interface 400. The window of the avatar may be pinned to any corner of the screen during the display.). Marey teaches: Claim 33. The method of claim 20, comprising:generating, using a machine learning model, an inference, based at least on input data that is based at least on the plurality of assets,wherein determining whether the scene comprises at least one character communicating using sign language is based at least on the inference (¶0047 Translation application server 104 includes a speech-to-text module 230, a text-to-sign language module 232, avatar generation module 234, or sentiment analysis module 236. Speech-to-text module 230 converts a speech received via microphone 226 of computing device 114 to text. Speech-to-text module 230 may implement any machine learning speech recognition or voice recognition techniques, such as Google® DeepMind, to decipher the speech of a user or a character in a content item. Text-to-sign language module 232 receives converted text generated by speech-to-text module 230 and translates the text to sign language (e.g., signs). The text can be translated to any sign language, such as American Sign Language, British Sign Language, or Spanish Sign Language. Text-to-sign language module 232 may utilize the sign language source 108 when translating to sign language.). Marey teaches: Claim 34. The method of claim 33, comprising translating the sign language into a language other than the sign language based at least on the inference (¶0047 Translation application server 104 includes a speech-to-text module 230, a text-to-sign language module 232, avatar generation module 234, or sentiment analysis module 236. Speech-to-text module 230 converts a speech received via microphone 226 of computing device 114 to text. Speech-to-text module 230 may implement any machine learning speech recognition or voice recognition techniques, such as Google® DeepMind, to decipher the speech of a user or a character in a content item. Text-to-sign language module 232 receives converted text generated by speech-to-text module 230 and translates the text to sign language (e.g., signs). The text can be translated to any sign language, such as American Sign Language, British Sign Language, or Spanish Sign Language. Text-to-sign language module 232 may utilize the sign language source 108 when translating to sign language.). Marey teaches: Claim 35. The method of claim 33, comprising generating, based on the plurality of game assets, the input data (¶0048 Avatar generation module 234 may generate a virtual avatar performing signs translated by text-to-sign language module 232. Avatar generation module 234 may use a Virtual Reality (VR) rendering technique, such as the Unreal® Engine, to generate a virtual avatar that mimics the user or character's movement and expression. Unreal® Engine is a software development environment suite for building virtual and augmented reality graphics, game development, architectural visualization, content creation, broadcast, or any other real-time applications). Marey teaches: Claim 36. The method of claim 33, comprising training, during a training phase, the machine learning model using labelled training data, the labelled training data comprising game assets and associated labels indicating whether the game assets represent scenes comprising one or more characters communicating using sign language (¶0047 Translation application server 104 includes a speech-to-text module 230, a text-to-sign language module 232, avatar generation module 234, or sentiment analysis module 236. Speech-to-text module 230 converts a speech received via microphone 226 of computing device 114 to text. Speech-to-text module 230 may implement any machine learning speech recognition or voice recognition techniques, such as Google® DeepMind, to decipher the speech of a user or a character in a content item. Text-to-sign language module 232 receives converted text generated by speech-to-text module 230 and translates the text to sign language (e.g., signs). The text can be translated to any sign language, such as American Sign Language, British Sign Language, or Spanish Sign Language. Text-to-sign language module 232 may utilize the sign language source 108 when translating to sign language.). As per claims 20, 37, 38, 40 and 41, the method, system and manufacture tracks the device of claims 1, 1, 23, 25 and 1, respectively, resulting in substantially similar limitations. The same cited prior art and rationale of claims 1, 1, 23, 25 and 1 are applied to claims 20, 37, 38, 40 and 41, respectively. Marey discloses that the embodiment may be found as a system and manufacture (Fig. 1 and ¶0038). Conclusion Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a). A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action. Any inquiry concerning this communication or earlier communications from the examiner should be directed to KURTIS GILLS whose telephone number is (571)270-3315. The examiner can normally be reached M-F 8-5 PM. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Jerry O’Connor can be reached on 571-272-6787. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /KURTIS GILLS/Primary Examiner, Art Unit 3624
Read full office action

Prosecution Timeline

Sep 23, 2024
Application Filed
Mar 10, 2026
Response after Non-Final Action
Apr 28, 2026
Non-Final Rejection mailed — §101, §103
Jul 27, 2026
Response Filed
Sep 17, 2026
Final Rejection mailed — §101, §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12749108
DETERMINING PRICING INFORMATION FROM MERCHANT DATA
2y 6m to grant Granted Sep 29, 2026
Patent 12737700
MATERIALS AND PROCESS INTEGRATION FOR BUILD PROJECTS WITH MOAB ASSEMBLIES
3y 1m to grant Granted Sep 15, 2026
Patent 12730675
AUTOMATION WITH COMPOSABLE ASYNCHRONOUS TASKS
3y 0m to grant Granted Sep 08, 2026
Patent 12731091
ANTI-MONOTONY SYSTEM AND METHOD ASSOCIATED WITH NEW HOME CONSTRUCTION IN A MASTER-PLANNED COMMUNITY
2y 1m to grant Granted Sep 08, 2026
Patent 12725117
AUXILIARY VERIFICATION SYSTEM OF GREENHOUSE GAS INVENTORY
2y 0m to grant Granted Sep 01, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

3-4
Expected OA Rounds
58%
Grant Probability
87%
With Interview (+29.2%)
3y 7m (~1y 6m remaining)
Median Time to Grant
Moderate
PTA Risk
Based on 565 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month