DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Prior arts cited in this office action:
Porikli et al. (US 20040239762 A1, hereinafter “Porikli”)
Van den Oord et al. (US 20180322891 A1 hereinafter “Van den Oord”)
Wang et al. (DE 102016013487 A1, hereinafter “Wang”)
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows:
1. Determining the scope and contents of the prior art.
2. Ascertaining the differences between the prior art and the claims at issue.
3. Resolving the level of ordinary skill in the pertinent art.
4. Considering objective evidence present in the application indicating obviousness or nonobviousness.
Claims 2-21 are rejected under 35 U.S.C. 103 as being unpatentable over Porikli et al. (US 20040239762 A1, hereinafter “Porikli”) in view Van den Oord et al. (US 20180322891 A1 hereinafter “Van den Oord”) and in view of Wang et al. (DE 102016013487 A1, hereinafter “Wang”).
Regarding claims 2, 10 and 18:
Porikli discloses a neural network system implemented by one or more computers, the neural network system being configured to receive a neural network input and to generate an output image from the neural network input, the output image comprising a plurality of pixels arranged in a two-dimensional map, each pixel having a respective color value for each of a plurality of color channels, and the neural network system comprising: one or more initial neural network layers configured to receive the neural network input and to process the neural network input to generate an alternative representation of the neural network input; and one or more output layers, wherein the output layers are configured to receive the alternative representation and to generate the output image pixel by pixel from a sequence of pixels taken from the output image, comprising, for each pixel in the output image, generating a respective score distribution over a discrete set of possible color values for each of the plurality of color channels (generation of new background images are construed by a neural network which considers individual pixels on a per color basis and taking color variance into account (Porikli [0052]-[0062], fig. 1 and 4).
Porikli fails to teach wherein processing the conditioning input using a self-attention-based encoder neural network to generate a sequential conditioning representation that comprises a sequence of encoded representations, wherein the self-attention-based encoder neural network neural network comprises one or more self-attention layers;
Van Denn Oord teaches the alternative representation from the convolutional subnetwork at each time step may be conditioned on a neural network input, for example a latent representation of a conditioning input. The conditioning input may be global (substantially time-independent) and/or local (time-dependent). The conditioning input may comprise, for example, text, image or video data, or audio data, for example an example of a particular speaker or language or music. The neural network input may comprise an embedding of the conditioning input. For example, in a text-to-speech system a global conditioning input may comprise a speaker embedding and a local conditioning input may comprise linguistic features. The system may be configured to map the neural network input, or a conditioning input, from a lower sampling frequency to the audio sample generation frequency, for example by repeating the input or upsampling the input using a neural network. Thus, the neural network input may comprise features of a text segment and the output sequence may represent a verbalization of the text segment; and/or the neural network input may comprise speaker or intonation pattern values; and/or the neural network input may include one or more of: speaker identity information, language identity information, and speaking style information. Alternatively, the output sequence represents a piece of music (Van Den Oord [0010]).
Wang further teaches in any case; the semantic attention system provides the attention weights described herein 906 which are used to control probabilistic classifications within the label generation model. At each iteration, a word is written in a sequence for the label using the attention weights 906 predicted to focus the model on certain concepts and attributes that are most relevant to that iteration. The attention weights 906 are reassessed and adjusted for each run (Wang [0069], fig. 10).
Therefore, taking the teaching of Porikli, Van den Oord as a whole, it would have been obvious to one of ordinary skill in the art before the effective filing date of the application to use self-attention encode based encoder neural network the generate the sequential conditioning representation, wherein each word is weight accordingly such that to help determine proper meaning, in order to enable high-speed training and captures complex, long-range dependencies across data, overcoming the bottlenecks and memory constraints of traditional sequential models. As a result, render the system more efficient.
Regarding claims 3, 11 and 21:
Porikli in view of Van den Oord and in view of Wang teaches wherein the second neural network comprises a sequence of subnetworks, one or more of the subnetworks comprising a respective encoder-decoder attention sub-layer that is configured to:
receive a current representation of the output image; and
update the current representation of the output image by applying an attention mechanism
over the encoded representations in the sequential conditioning representation using one or more
queries derived from the current representation of the output image (Van den Oord [0011], [0065], figs. 1 and 3; Wang [0024], [0029]).
Regarding claims 4, 12 and 19:
Porikli in view of Van den Oord and in view of Wang teaches wherein the conditioning input is a text sequence that describes the output image (Van den Oord [0010], [0013] [0065], figs. 1 and 3; Wang [0024], [0029]).
Regarding claims 5, 13 and 20:
Porikli in view of Van den Oord and in view of Wang teaches wherein the conditioning input is another image (Van den Oord [0010], [0013] [0065], figs. 1 and 3; Wang [0024], [0029]).
Regarding claims 6 and 14:
Porikli in view of Van den Oord and in view of Wang teaches wherein the encoder neural network comprises one or more un-masked self-attention layers (Van den Oord [0011], [0065], figs. 1 and 3; Wang [0024], [0029]).
Regarding claims 7 and 15:
Porikli in view of Van den Oord and in view of Wang teaches wherein one or more of the subnetworks comprise a respective self-attention layer (Van den Oord [0011], [0065], figs. 1 and 3; Wang [0024], [0029]).
Allowable Subject Matter
Claims 8-9, 16 and 17 are objected to as being dependent upon a rejected base claim but would be allowable if rewritten in independent form including all of the limitations of the base claim and any intervening claims.
Conclusion
Any inquiry concerning this communication or earlier communications from the examiner should be directed to WEDNEL CADEAU whose telephone number is (571)270-7843. The examiner can normally be reached Mon-Fri 9:00-5:00.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Chieh Fan can be reached at 571-272-3042. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/WEDNEL CADEAU/Primary Examiner, Art Unit 2632 August 12, 2026