DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Specification
The title of the invention is not descriptive. A new title is required that is clearly indicative of the invention to which the claims are directed.
The following title is suggested: -- Reduced Data Transmission and Synthesis Method and System using Artificial Intelligence--.
Claim Objections
Claim 5 is objected to because of the following informalities:
In Claim 5, "the data transmission method further comprise" should be corrected to read --the data transmission method further comprises--.
Appropriate correction is required.
Examiner Notes on Patent Subject Matter Eligibility under 35 U.S.C. 101
Independent Claims 1, 7, 13, and 17 are directed towards various embodiments for abstract data transformation of voice or video data for network transmission using an artificial intelligence model. While certain limitations such as generating voice or video data could be performed by a human by speaking or moving/making facial expressions, the claims all involve a combination of AI model processing for data abstraction coupled with network transmission/reception of such data that uses reduce a data rate for voice/video calling (see Specification, Paragraph 0051). Accordingly, these claims and their dependents are directed towards a practical application under step 2A prong 2 of the 2019 Patent Subject Matter Eligibility Guidelines and are directed towards patent eligible subject matter under 35 U.S.C. 101.
Claim Rejections - 35 USC § 112
The following is a quotation of 35 U.S.C. 112(b):
(b) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention.
The following is a quotation of 35 U.S.C. 112 (pre-AIA ), second paragraph:
The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the applicant regards as his invention.
Claims 7-20 are rejected under 35 U.S.C. 112(b) or 35 U.S.C. 112 (pre-AIA ), second paragraph, as being indefinite for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor (or for applications subject to pre-AIA 35 U.S.C. 112, the applicant), regards as the invention.
Claim 7 recites "the transmitting device" in line 3 and again in line 5. There is insufficient antecedent basis for this term and it is unclear whether the term should refer back to the "transmitting apparatus" introduced in line 2 or is referencing some additional device. For claim construction in the interest of compact prosecution, "the transmitting device" will be construed as --the transmitting apparatus--. Also, "the receiving device" found in line 13 similarly lacks antecedent basis and it is unclear whether the "receiving apparatus" or another receiving device is intended. For claim construction "the receiving device" will be construed as --the receiving apparatus--.
In Claim 11, "the data transmission method" lacks antecedent basis and it is unclear what prior term is being referenced. For claim interpretation this limitation the remainder of the claimed synthesis will be construed as being introduced by a wherein clause (e.g., --wherein the receiving apparatus synthesizes the synthesized voice data...--).
Claim 13 recites "the transmitting apparatus" in line 4. There is insufficient antecedent basis for this term and it is unclear whether the term should refer back to the "transmitting device" introduced in line 3 or is referencing some additional device. For claim construction in the interest of compact prosecution, "the transmitting apparatus" will be construed as --the transmitting device--.
Claim 17 recites the transmitting device" in line 6. There is insufficient antecedent basis for this term and it is unclear whether the term should refer back to the "transmitting apparatus" introduced in lines 2-3 or is referencing some additional device. For claim construction in the interest of compact prosecution, "the transmitting device" will be construed as --the transmitting apparatus--.
The dependent claims inherit and fail to resolve the indefinite claim language of their parent claim(s), and thus, are also indefinite under 35 U.S.C. 112(b) by virtue of their dependency.
Claim Rejections - 35 USC § 102
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action:
A person shall be entitled to a patent unless –
(a)(2) the claimed invention was described in a patent issued under section 151, or in an application for patent published or deemed published under section 122(b), in which the patent or application, as the case may be, names another inventor and was effectively filed before the effective filing date of the claimed invention.
Claims 1-3, 6-9, 12-14, and 17-18 are rejected under 35 U.S.C. 102(a)(2) as being anticipated by Mahapatra, et al. (U.S. Patent: 12,126,791).
With respect to Claim 1, Mahapatra discloses:
A data transmission method, comprising:
generating, by a transmitting apparatus of a data transmission system, a voice data for a voice call or a video data for a video call (videoconferencing communication system including a client device that generates both video and voice audio content, Col. 5, Lines 35-42; Col. 6, Lines 7-32; Col. 15, Lines 42-57; and Col. 16, Lines 24-57);
transforming, by the transmitting apparatus, the voice data or the video data into an abstract data according to a first artificial intelligence (AI) model, wherein the abstract data comprises information related to the voice data or the video data, and a size of the abstract data is smaller than a size of the voice data or the video data (conversion of video content using neural network models to generate an abstraction of the video data, Col. 10, Lines 50-62 and Col. 12, Line 13- Col. 13, Line 51, where the processing "provides a substantial reduction in video size" that leads to "a significant reduction in bandwidth," Col. 3, Lines 43-57);
transmitting, by the transmitting apparatus, the abstract data to a receiving apparatus of the data transmission system (transmission of the video representation rather than the full video file by a transmitting device to a receiving device, Abstract; Col. 7, Lines 20-64; Col. 16, Line 58- Col. 17, Line 3; Col. 19, Line 65- Col. 20, Line 20; see also Fig. 2A showing multiple computing devices for multiple participants);
synthesizing, by the receiving apparatus, a synthesized voice data or a synthesized video data from the abstract data according to a second AI model (receiving participant device uses a video synthesizer neural network to synthesize video data using a neural network along with audio data, Col. 5, Lines 33-42; Col. 17, Lines 36-54; Col. 18, Line 38- Col. 19, Line 14; and Col. 20, Lines 2-10); and
playing, by the receiving apparatus, the synthesized voice data or the synthesized video data (synthesized video and audio voice playback, Col. 5, Lines 33-42; Col. 15, Lines 29-31; Col. 18, Lines 38-55; and Col. 26, Lines 47-49).
With respect to Claim 2, Mahapatra further discloses:
The data transmission method of claim 1, wherein the information of the abstract data comprises media description, words, phrases, and emotional information in the voice data or the video data (emotion information is represented for the video abstraction such as facial expressions or behaviors using the neural network models, Col. 5, Lines 45-54; Col. 10, Line 23- Col. 11, Line 63; and Col. 13, Lines 35-51).
With respect to Claim 3, Mahapatra further discloses:
The data transmission method of claim 1, wherein the first AI model comprises at least one of a hidden Markov model (HMM) model and a neural network model (HMM and/or neural network, Col. 12, Lines 15-26; note that the language "at least one of...and" is construed as being disjunctive since the specification does not require all embodiments to use both models together where the models are disclosed in the alternative as per paragraph 0052 and 0101 (note: "may comprise")).
With respect to Claim 6, Mahapatra further discloses:
The data transmission method of claim 1, wherein the information of the abstract data further comprises image information in an event that the abstract data is generated based on the video data (image information representative of the input video data, Col. 13, Line 1- Col. 14, Line 23; Col. 15, Lines 48-57; and Abstract).
With respect to Claim 7, Mahapatra discloses:
A data transmission system, comprising:
a transmitting apparatus (computer device having a communication module for transmission, Col. 6, Line 56- Col. 7, Line 64; Abstract; Fig. 2B);
a network node (network having computer nodes, Col. 7, Lines 20-64 and Col. 28, Lines 5-32); and
a receiving apparatus, wirelessly communicating with the transmitting device through the network node (computer device having a communication module for reception of conference communications over a network node, Col. 6, Line 56- Col. 7, Line 64; Abstract; Fig. 2B),
wherein the transmitting device generates a voice data for a voice call or a video data for a video call (videoconferencing communication system including a client device that generates both video and voice audio content, Col. 5, Lines 35-42; Col. 6, Lines 7-32; Col. 15, Lines 42-57; and Col. 16, Lines 24-57), transforms the voice data or the video data into an abstract data according to a first artificial intelligence (AI) model, wherein the abstract data comprises information related to the voice data or the video data, and a size of the abstract data is smaller than a size of the voice data or the video data (conversion of video content using neural network models to generate an abstraction of the video data, Col. 10, Lines 50-62 and Col. 12, Line 13- Col. 13, Line 51, where the processing "provides a substantial reduction in video size" that leads to "a significant reduction in bandwidth," Col. 3, Lines 43-57), and transmits the abstract data to the receiving apparatus (transmission of the video representation rather than the full video file by a transmitting device to a receiving device, Abstract; Col. 7, Lines 20-64; Col. 16, Line 58- Col. 17, Line 3; Col. 19, Line 65- Col. 20, Line 20; see also Fig. 2A showing multiple computing devices for multiple participants), and
wherein the receiving device synthesizes a synthesized voice data or a synthesized video data from the abstract data according to a second AI model (receiving participant device uses a video synthesizer neural network to synthesize video data using a neural network along with audio data, Col. 5, Lines 33-42; Col. 17, Lines 36-54; Col. 18, Line 38- Col. 19, Line 14; and Col. 20, Lines 2-10, and plays the synthesized voice data or the synthesized video data (synthesized video and audio voice playback, Col. 5, Lines 33-42; Col. 15, Lines 29-31; Col. 18, Lines 38-55; and Col. 26, Lines 47-49).
Claims 8-9 and 12 contain subject matter respectively similar to Claims 2-3 and 6, and thus, are rejected under similar rationale.
With respect to Claim 13, Mahapatra discloses:
A data transmission method, comprising:
performing, by a processor of a receiving apparatus, a voice call or a video call with a transmitting device (videoconferencing communication system including a client device that generates both video and voice audio content that can receive voice or video call communications from a transmitting device wherein a receiving computing device includes a processor for carrying out the disclosed process steps, Col. 5, Lines 35-42; Col. 6, Lines 7-32; Col. 6, Line 56- Col. 7, Line 64; Col. 15, Lines 42-57; and Col. 16, Lines 24-57; Fig. 2B showing multiple computing devices for multiple participants);
receiving, by the processor, an abstract data from the transmitting apparatus, wherein the abstract data comprises information related to a voice data for the voice call or a video data for the video call, and a size of the abstract data is smaller than a size of the voice data or the video data (reception of a video representation rather than the full video file from a transmitting device by a receiving device, Abstract; Col. 7, Lines 20-64; Col. 16, Line 58- Col. 17, Line 3; Col. 19, Line 65- Col. 20, Line 20; see also Fig. 2A showing multiple computing devices for multiple participants);
synthesizing, by the processor, a synthesized voice data or a synthesized video data from the abstract data according to an artificial intelligence (AI) model (receiving computing device uses a video synthesizer neural network to synthesize video data using a neural network along with audio data, Col. 5, Lines 33-42; Col. 17, Lines 36-54; Col. 18, Line 38- Col. 19, Line 14; and Col. 20, Lines 2-10); and
playing, by the processor, the synthesized voice data or the synthesized video data (synthesized video and audio voice playback, Col. 5, Lines 33-42; Col. 15, Lines 29-31; Col. 18, Lines 38-55; and Col. 26, Lines 47-49).
Claim 14 contains subject matter similar to Claims 2 and 6, and thus, is rejected under similar rationale.
With respect to Claim 17, Mahapatra discloses:
An apparatus, comprising:
a transceiver which, during operation, wirelessly communicates with a transmitting apparatus through a network node (computer device having a communication module for transmission/receiving of communication data over a network node during an audiovisual conference, Col. 6, Line 56- Col. 7, Line 64; Col. 28, Lines 5-32; Col. 8, Lines 20-21 (see- “receive and send”); Abstract; Fig. 2B); and
a processor communicatively coupled to the transceiver such that, during operation, the processor performs operations comprising (computing device processor in communication/connection with the communication module (see discussion of a bus and control over sending and receiving) that implements program instructions, Col. 6, Line 56- Col. 7, Line 64; Col. 8, Lines 20-21):
performing a voice call or a video call with the transmitting device (videoconferencing participation wherein a client device receives voice or video call communications from a transmitting device wherein a receiving computing device includes a processor for carrying out the disclosed process steps, Col. 5, Lines 35-42; Col. 6, Lines 7-32; Col. 6, Line 56- Col. 7, Line 64; Col. 15, Lines 42-57; and Col. 16, Lines 24-57; Fig. 2B showing multiple computing devices for multiple participants);
receiving, via the transceiver, an abstract data from the transmitting apparatus, wherein the abstract data comprises information related to a voice data for the voice call or a video data for the video call, and a size of the abstract data is smaller than a size of the voice data or the video data (reception of a video representation with the communication module rather than the full video file from a transmitting device by a receiving device, Abstract; Col. 7, Lines 20-64; Col. 16, Line 58- Col. 17, Line 3; Col. 19, Line 65- Col. 20, Line 20; see also Fig. 2A showing multiple computing devices for multiple participants);
synthesizing a synthesized voice data or a synthesized video data from the abstract data according to an artificial intelligence (AI) model (receiving computing device uses a video synthesizer neural network to synthesize video data using a neural network along with audio data, Col. 5, Lines 33-42; Col. 17, Lines 36-54; Col. 18, Line 38- Col. 19, Line 14; and Col. 20, Lines 2-10); and
playing the synthesized voice data or the synthesized video data (synthesized video and audio voice playback, Col. 5, Lines 33-42; Col. 15, Lines 29-31; Col. 18, Lines 38-55; and Col. 26, Lines 47-49).
Claim 18 contains subject matter similar to Claims 2 and 6, and thus, is rejected under similar rationale.
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claims 4, 10, 15, and 19 are rejected under 35 U.S.C. 103 as being unpatentable over Mahapatra, et al. in view of MacConnell, et al. (U.S. PG Publication: 2020/0211540 A1).
With respect to Claim 4, Mahapatra teaches the method for compressed data transmission including audio and video data in a conferencing platform. Although Mahapatra teaches the transmission of a transcription of voice for rendering at a recipient (Col. 5, Lines 33-42 and Col. 16, Line 58- Col. 17, Line 3) and speech synthesis (Col. 15, Lines 25-39), Mahapatra does not disclose that the second AI model comprises a text-to-speech (TTS) model. MacConnell, however, discloses that speech audio in a communication is converted and transmitted as text where it is converted back to speech audio using a neural network text-to-speech synthesis model (Paragraphs 0018 and 0023-0027).
Mahapatra and MacConnell are analogous art because they are from a similar field of endeavor in the field of communication content reduction using machine learning. Thus, it would have been obvious to one of ordinary skill in the art, before the effective filing date, to utilize the text-to-speech model taught by MacConnell to synthesize the conference call transcription taught by Mahapatra to provide a model that can learn to accurately synthesize speech to further conserve communication bandwidth (MacConnell, Paragraph 0018).
Claims 10, 15, and 19 contain subject matter similar to Claim 4, and thus, are rejected under similar rationale.
Claims 5, 11, 16, and 20 are rejected under 35 U.S.C. 103 as being unpatentable over Mahapatra, et al. in view of Liu, et al. (U.S. PG Publication: 2023/0035306 A1).
With respect to Claim 5, Mahapatra teaches the method for compressed data transmission including audio and video data in a conferencing platform. Although Mahapatra teaches
the synthesis of video data and rendering of audio data (Col. 5, Lines 33-42; Col. 15, Lines 18-41; Col. 16, Line 58- Col. 17, Line 3; Col. 18, Lines 38-55), Mahapatra does not teach voice print information is stored in the receiving apparatus and its use in synthesis. Liu, however, discloses a recipient storing "voice reference data" (i.e., a voice print) that is used by a text-to-speech neural network at a receiving device (Paragraphs 0062 and 0065; Fig. 2B, Element 254).
Mahapatra and Liu are analogous art because they are from a similar field of endeavor in communication content reduction using machine learning. Thus, it would have been obvious to one of ordinary skill in the art, before the effective filing date, to utilize the voice print information taught by Liu in the conference audio synthesis/playback of Liu in order to better customize output speech and further reduce communication content for reproduction (Liu, Paragraph 0048).
Claims 11, 16, and 20 contain subject matter similar to Claim 5, and thus, are rejected under similar rationale.
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure:
Giovanardi, et al. (U.S. PG Publication: 2024/0364653 A1)- teaches a trained machine learning model that summarizes unread messages from a chat and video conference (Paragraph 0082).
Any inquiry concerning this communication or earlier communications from the examiner should be directed to JAMES S WOZNIAK whose telephone number is (571)272-7632. The examiner can normally be reached 7-3, off alternate Fridays.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant may use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Andrew Flanders can be reached at (571)272-7516. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
JAMES S. WOZNIAK
Primary Examiner
Art Unit 2655
/JAMES S WOZNIAK/Primary Examiner, Art Unit 2655