DETAILED ACTION
1. This action is responsive to Application no.19/038,118 filed 1/27/2025. All claims have been examined and are currently pending.
Notice of Pre-AIA or AIA Status
2. The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Information Disclosure Statement
3. The information disclosure statement (IDS) submitted is in compliance with the provisions of 37 CFR 1.97. Accordingly, the information disclosure statement is being considered by the examiner.
Claim Rejections - 35 USC § 102
4. In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
5. The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action:
A person shall be entitled to a patent unless –
(a)(1) the claimed invention was patented, described in a printed publication, or in public use, on sale, or otherwise available to the public before the effective filing date of the claimed invention.
(a)(2) the claimed invention was described in a patent issued under section 151, or in an application for patent published or deemed published under section 122(b), in which the patent or application, as the case may be, names another inventor and was effectively filed before the effective filing date of the claimed invention.
6. Claims 1-20 are rejected under 35 U.S.C. 102(a)(1) as being anticipated by Peralta et al (2023/0316952).
Regarding claim 1 Peralta et al (2023/0316952) teaches One or more processors comprising processing circuitry (figure 1, 9: communication capable device, 10 CPU; para 0013: system and method for bidirectional automatic sign language translation and visualization; communication-capable devices; 96-97 computer system) to:
receive an inbound data feed comprising communications content data (0030: audio sensor 402 that is able to capture sounds; interface 403 that is able to capture textual information; 0042);
obtain a sign language preference associated with an output data feed (0049: where the display or presentation is according to request output type 401; [0050] The method where the input of production block 400 is a feed of spoken language, and optional information regarding what media or type the output is expected to be, herein referred to as request output type 401. The expected output of production block 400 is the conversion from spoken language to sign language.;
[0054] According to FIG. 4 of one or more embodiments of the present invention, sign language identification module 500 enables communication-capable devices (e.g., user device) 200 to take note of or save the following information including any information related to it, such as which languages the user typically signs in and which language the user typically requests translations for. Sign language input 501 is acquired and subsequently processed in sign language identification module 502. Processed data is then further inputted to production method block 400 and translation method block 300.);
generate sign language translation data comprising a translation of the communications content data based at least on the sign language preference (0043: sign language translation; 44-49; 0049; 0050; 0054);
generate sign language video data representing a visual representation of the sign language translation data ([0049] (f) An output processing module 408 adjusts or packages sign language poses to create output data 409 that may be displayed or presented in the device (e.g., communication-capable devices 200a and 200b) of the sign language users, where the display or presentation is according to requested output type 401.); and
cause a user interface to present the visual representation of the sign language translation data (fig 1; 0049: create output data 409 that may be displayed or presented in the device (e.g., communication-capable devices 200a and 200b) of the sign language users).
Regarding claim 2 Peralta teaches The one or more processors of claim 1, wherein the communications content data comprises at least one of spoken word audio content data, text content data, video content data, or incoming sign language content data (0030: spoken; written; 0044: audio feed).
Regarding claim 3 Peralta teaches The one or more processors of claim 1, wherein the processing circuitry is further to control a multi-user virtual environment to render one or more avatars within the multi-user virtual environment to present the communications content data based at least on the sign language video data (0049; 0069-0071: output, 3D avatar animation; 85-87).
Regarding claim 4 Peralta teaches The one or more processors of claim 1, wherein the processing circuitry is further to generate the sign language translation data based at least on a sign language determined based at least on the sign language preference (0049: where the display or presentation is according to request output type 401; [0050] The method where the input of production block 400 is a feed of spoken language, and optional information regarding what media or type the output is expected to be, herein referred to as request output type 401. The expected output of production block 400 is the conversion from spoken language to sign language.;
[0054] According to FIG. 4 of one or more embodiments of the present invention, sign language identification module 500 enables communication-capable devices (e.g., user device) 200 to take note of or save the following information including any information related to it, such as which languages the user typically signs in and which language the user typically requests translations for. Sign language input 501 is acquired and subsequently processed in sign language identification module 502. Processed data is then further inputted to production method block 400 and translation method block 300.).
Regarding claim 5 Peralta teaches The one or more processors of claim 1, wherein the processing circuitry is further to obtain the sign language preference based on reading a preference setting of a client application that generates the user interface
([0054] According to FIG. 4 of one or more embodiments of the present invention, sign language identification module 500 enables communication-capable devices (e.g., user device) 200 to take note of or save the following information including any information related to it, such as which languages the user typically signs in and which language the user typically requests translations for. Sign language input 501 is acquired and subsequently processed in sign language identification module 502. Processed data is then further inputted to production method block 400 and translation method block 300).
Regarding claim 6 Peralta teaches The one or more processors of claim 1, wherein the processing circuitry is further to obtain the sign language preference based at least on a detection of a form of sign language from video data of a user of the user interface
([0054] According to FIG. 4 of one or more embodiments of the present invention, sign language identification module 500 enables communication-capable devices (e.g., user device) 200 to take note of or save the following information including any information related to it, such as which languages the user typically signs in and which language the user typically requests translations for. Sign language input 501 is acquired and subsequently processed in sign language identification module 502. Processed data is then further inputted to production method block 400 and translation method block 300.).
Regarding claim 7 Peralta teaches The one or more processors of claim 6, wherein the processing circuitry is further to determine the form of sign language based on applying the video data to one or more machine learning models to infer the form of sign language (0031-0037 translation, process the input data by using Deep Neural Network (DNN) models; 54-55 sign language identification; machine learning).
Regarding claim 8 Peralta teaches The one or more processors of claim 1, wherein the processing circuitry is further to generate the sign language video data to comprises at least a portion of an animated avatar, wherein the at least the portion of the animated avatar performs signing corresponding to the sign language translation data (0049; 0069-0071: output, 3D avatar animation; 85-87).
Regarding claim 9 Peralta teaches The one or more processors of claim 1, wherein the processing circuitry is further to generate the sign language video data using a sign language version indicated by the sign language preference, based at least on an augmentation of an appearance of a presenter of the communications content data (0049; 0051: photorealistic video; 0054; 0069-0071: output, photo-realistic; 85-87).
Regarding claim 10 Peralta teaches The one or more processors of claim 1, wherein the one or more processors are further to execute a framework comprising one or more machine learning models that generate the sign language translation data based at least on the communications content data and a sign language indicated by the sign language preference (0031-0037 translation; process the input data by using Deep Neural Network (DNN) models; 54-55 sign language identification; machine learning).
Regarding claim 11 Peralta teaches The one or more processors of claim 1, wherein the one or more processors are further to execute a framework comprising one or more machine learning models that generate the sign language video data representing the visual representation of the sign language translation data based at least on the sign language translation data (43-49: production; sign language translation; 55: production and translation; machine learning).
Regarding claim 12 Peralta teaches The one or more processors of claim 1, wherein the processing circuitry is comprised in at least one of:
a control system for an autonomous or semi-autonomous machine;
a perception system for an autonomous or semi-autonomous machine;
a system for performing simulation operations;
a system for performing digital twin operations;
a system for performing light transport simulation;
a system for performing collaborative content creation for three-dimensional assets;
a system for performing deep learning operations;
a system for performing remote operations;
a system for performing real-time streaming;
a system for generating or presenting one or more of augmented reality content, virtual reality content, or mixed reality content;
a system implemented using an edge device;
a system for generating or presenting virtual reality (VR) content;
a system for generating or presenting augmented reality (AR) content;
a system for generating or presenting mixed reality (MR) content;
a system implemented using a robot;
a system for performing conversational Al operations;
a system implementing one or more language models;
a system implementing one or more large language models (LLMs);
a system implementing one or more small language models (SLMs);
a system implementing one or more vision language models (VLMs);
a system implementing one or more multimodal language models (MMLMs);
a system implemented using one or more cloud-hosted microservices;
a system for generating synthetic data;
a system for generating synthetic data using Al;
a system incorporating one or more virtual machines (VMs);
a system implemented at least partially in a data center; or
a system implemented at least partially using cloud computing resources (96: may be a cloud computing).
Regarding claim 13 Peralta teaches A system comprising one or more processors (13: system; 96-97: processor) to:
receive a first data feed comprising communications content data (0030: spoken; written; 0042);
obtain a sign language preference associated with a second data feed (49-50; 0054: which languages the user typically signs in; sign language input is acquired and subsequently processed in sign language identification module);
generate sign language translation data comprising a translation of the communications content data based at least on a language indicated by the sign language preference ([0042] FIG. 3 details the production method block 400 according to one or more embodiments. Audio feed 402a is acquired from audio sensor 402. Audio feed is converted to text, patterns are identified in the text, and the data obtained is utilized to train DNN models to compile a sequence of poses and thus generate output media, such as photorealistic and/or animated videos by output processor module 408.;
43-50; 54); and
cause a user interface to output an augmented representation of the first data feed based on the sign language translation data ([0049] (f) An output processing module 408 adjusts or packages sign language poses to create output data 409 that may be displayed or presented in the device (e.g., communication-capable devices 200a and 200b) of the sign language users, where the display or presentation is according to requested output type 401).
Regarding claim 14 Peralta teaches The system of claim 13, the one or more processors further to generate a visual representation of the sign language translation data, wherein the representation of the sign language translation data comprises the visual representation of the sign language translation data (49).
Regarding claim 15 Peralta teaches The system of claim 14, the one or more processors further to generate the visual representation of the sign language translation data to comprise at least a portion of an animated avatar, wherein the portion of the animated avatar performs signing corresponding to the sign language translation data using a sign language version indicated by the sign language preference (0049; 0069-0071: output, 3D avatar animation; 85-87).
Regarding claim 16 Peralta teaches The system of claim 14, wherein the one or more processors are further to generate the visual representation of the sign language translation data based at least on an augmentation of an appearance of a presenter of the communications content data as represented by the first data feed (0049; 0051: photorealistic video; 0054; 0069-0071: output, photo-realistic; 85-87).
Regarding claim 17 Peralta teaches The system of claim 13, the one or more processors further to generate an audio representation of the sign language translation data, wherein the representation of the sign language translation data comprises the audio representation of the sign language translation data (31: translation method…converts visual data in sign language into spoken language as output).
Regarding claim 18 Peralta teaches The system of claim 13, the one or more processors further to execute a framework comprising one or more machine learning models that generate the sign language translation data based at least on the first data feed comprising the communications content data and a sign language version indicated by the sign language preference (0031-0037 translation; process the input data by using Deep Neural Network (DNN) models; 54-55 sign language identification; machine learning).
Regarding claim 19 Peralta teaches The system of claim 13, wherein the system is comprised in at least one of:
a control system for an autonomous or semi-autonomous machine;
a perception system for an autonomous or semi-autonomous machine;
a system for performing simulation operations;
a system for performing digital twin operations;
a system for performing light transport simulation;
a system for performing collaborative content creation for three-dimensional assets;
a system for performing deep learning operations;
a system for performing remote operations;
a system for performing real-time streaming;
a system for generating or presenting one or more of augmented reality content, virtual reality content, or mixed reality content;
a system implemented using an edge device;
a system for generating or presenting virtual reality (VR) content;
a system for generating or presenting augmented reality (AR) content;
a system for generating or presenting mixed reality (MR) content;
a system implemented using a robot;
a system for performing conversational Al operations;
a system implementing one or more language models;
a system implementing one or more large language models (LLMs);
a system implementing one or more small language models (SLMs);
a system implementing one or more vision language models (VLMs);
a system implementing one or more multimodal language models (MMLMs);
a system implemented using one or more cloud-hosted microservices;
a system for generating synthetic data;
a system for generating synthetic data using Al;
a system incorporating one or more virtual machines (VMs);
a system implemented at least partially in a data center; or
a system implemented at least partially using cloud computing resources (96: may be a cloud computing).
Regarding claim 20 Peralta teaches A method (0013 method) comprising:
generating sign language translation data comprising a translation of downlink content data based at least on a determination of a sign language preference associated with uplink content data, wherein the sign language preference indicates a form of sign language (0013; 0042-0043; 0054); and
controlling a user interface to output a representation of the sign language translation data, wherein the representation of the sign language translation data comprises at least one of a visual representation of the sign language translation data or an audio representation of the sign language translation data (0013; 0026; 0038-0039; 0042).
Recites limitations similar to claims 1 and 13 and is rejected for similar rationale and reasoning
Conclusion
7. The prior art made of record and not relied upon is considered pertinent to applicant's disclosure: See PTO-892.
Gilbert et al 2009/0012788
Abstract: The translation system of a preferred embodiment includes an input element that receives an input language as audio information, an output element that displays an output language as visual information, and a remote server coupled to the input element and the output element, the remote server including a database of sign language images; and a processor that receives the input language from the input element, translates the input language into the output language, and transmits the output language to the output element, wherein the output language is a series of the sign language images that correspond to the input language and that are coupled to one another with substantially seamless continuity, such that the ending position of a first image is blended into the starting position of a second image.
Mahgoub et al (2026/00304559)
Abstract: The present invention facilitates communication between sign language users and machines by translating sign language and text using AI models, deep learning computer vision, and word embeddings. Users interact via sign language, captured and processed through deep learning and NLP modules. The system converts sign language videos into text, constructs coherent sentences, and generates contextually appropriate responses
Zargari-Khuzani et al (2025/0218223)
Abstract: System and techniques to facilitate the translation of a sign language into another language are described herein. A modular architecture may be used in which the output of different classifiers may be used to produce intermediate representations, or final translations, of the sign language. These classifiers may be trained on different types of signs to enhance accuracy while reduce training time and complexity.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to SHAUN A ROBERTS whose telephone number is (571)270-7541. The examiner can normally be reached Monday-Friday 9-5 EST.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Andrew Flanders can be reached on 571-272-7516. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov.
For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative or access to the automated information system, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/SHAUN ROBERTS/Primary Examiner, Art Unit 2655