Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
DETAILED ACTION
Response to Arguments
Applicant’s arguments, see pgs. 9-10 of applicant’s arguments, filed 06/29/2026, with respect to the rejection(s) of claim(s) 1-9 under 103 have been fully considered and are persuasive. Therefore, the rejection has been withdrawn. However, upon further consideration, a new ground(s) of rejection is made in view of Liew. Furthermore, new claims 10-13 are also rejected.
Claim Rejections - 35 USC § 102
The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action:
A person shall be entitled to a patent unless –
(a)(1) the claimed invention was patented, described in a printed publication, or in public use, on sale, or otherwise available to the public before the effective filing date of the claimed invention.
(a)(2) the claimed invention was described in a patent issued under section 151, or in an application for patent published or deemed published under section 122(b), in which the patent or application, as the case may be, names another inventor and was effectively filed before the effective filing date of the claimed invention.
Claims 1-5, 8-9, 13 are rejected under 35 U.S.C. 102(a)(1) as being anticipated by Liew (Pub No. US 20240144544 A1).
As per claim 1, Liew anticipates the claimed:
1. 1. (Currently Amended) An image generation apparatus comprising: at least one memory configured to store program code;
And at least one processor configured to operate as instructed by the program code, the program code including: acquisition code configured to cause at least one of the at least one processor to acquire text information and a prompt, (Liew [0028]: “It is to be understood that a text prompt may be a conditioner for a text-to-image model to generate an image that is semantically consistent with the text prompt e.g., by optimizing the latent vector or the generator to maximize the similarity between the text prompt and the image. That is, the text-to-image model may generate an image conditioned on or consistent with the conditioner (e.g., a text prompt).” The conditioner is the prompt that is separate from the text information.).
the prompt being an instruction, that is separate from the text information, to output metadata that includes a plurality of attribute values composed of attribute values that respectively correspond to a plurality of attributes and semantically match the text information; (Liew [0028]: “It is to be understood that a text prompt may be a conditioner for a text-to-image model to generate an image that is semantically consistent with the text prompt e.g., by optimizing the latent vector or the generator to maximize the similarity between the text prompt and the image. That is, the text-to-image model may generate an image conditioned on or consistent with the conditioner (e.g., a text prompt).”
Liew [0026]: “It is to be understood that “pre-trained text-to-image diffusion-based generative model” may refer to a pre-trained (described above), diffusion-based (having a diffusing process and a de-noising process from a diffusion model, described above), text-to-image generative model (described above). In an example embodiment, a text-to-image diffusion-based generative models may refer to a diffusion-based generative model that accepts a text input and synthesizes an image matching the text input. It will be appreciated that a machine learning model, such as a text-to-image diffusion-based generative model, may transform an input text into a latent representation to produce an image conditioned on that latent representation.” The latent representation is the instruction that is separate from the text information.).
derivation code configured to cause at least one of the at least one processor to derive metadata corresponding to the text information by inputting the text information and the prompt to a language model configured to infer metadata that semantically matches the description of input text, and output metadata in accordance with instructions of a prompt; (Liew [0028]: “It is to be understood that a text prompt may be a conditioner for a text-to-image model to generate an image that is semantically consistent with the text prompt e.g., by optimizing the latent vector or the generator to maximize the similarity between the text prompt and the image. That is, the text-to-image model may generate an image conditioned on or consistent with the conditioner (e.g., a text prompt).”).
and generation code configured to cause at least one of the at least one processor to generate an image based on the derived metadata. (Liew [0001]: “The embodiments described herein pertain generally to generating an object using a diffusion model. More specifically, the embodiments described herein pertain to generating an object from text and/or image input using a pre-trained text-to-image diffusion-based generative model.”).
As per claims 16 and 17, these claims are similar in scope to limitations recited in claim 3, and thus are rejected under the same rationale.
As per claim 2, Liew anticipates the claimed:
2. The image generation apparatus according to claim 1, further comprising a storage configured to store a plurality of image parts, wherein the generation code is configured to cause at least one of the at least one processor to select a predetermined number of image parts from the storage based on the plurality of attribute values included in the derived metadata, and generate the image by combining the predetermined number of image parts. (Liew fig. 6 shows different image parts with associate metadata. Liew [0031]: “As referenced herein, “similarity” may refer to a numeric value representing a degree of how close two objects (e.g., two images, two concepts corresponding to respective objects, etc.) are when the two objects are compared. It is to be understood that a similarity between two objects may be determined by using e.g., technologies such as sum of squared differences, mutual information, normalized mutual information, cross-correlation, etc. In an example embodiment, the higher the similarity (or value), the more contextually similar the two objects are. In such embodiment, a similarity (or value) “0” may indicate that the two objects are completely different. It is to be understood that in another example embodiment, the lower the similarity (or value), the more contextually similar the two objects are. In such embodiment, a similarity (or value) “0” may indicate that the two objects are identical.” The similarity values relate to the attribute values. Liew [0055]: “It is to be understood that a desired injection time (or a desired k, the desired number of operations or iterations) may be determined based on the determined semantic similarity value N between the two concepts. For example, as shown in the upper row of the images in FIG. 5A, when mixing or blending “rabbit” and “hamster” (two similar concepts or objects), the model may require fewer de-noising operations or iterations to generate features of the hamster due to the high level of similarity between a rabbit and a hamster (e.g., only need to modify the eyes and nose of the rabbit). As such, the number k may be adjusted (e.g., decreased) based on the determined similarity value N. In another example, as shown in the lower row of the images in FIG. 5A, when mixing or blending “rabbit” and “coffee machine” (two extremely dissimilar concepts or objects), the model may require more de-noising operations or iterations to generate the features of the coffee machine and/or to overwrite the rabbit details due to their high level of dissimilarity. As such, the number k may be adjusted (e.g., increased) based on the determined similarity value N.” These different features can be combined or blended, and the model accounts for the combination of dissimilar concepts.).
As per claim 3, Liew anticipates the claimed:
3. The image generation apparatus according to The image generation apparatus according to wherein the plurality of image parts stored in the storage are each provided with a name composed of one or more character strings that express the image part, (Liew [0002]: “A text-to-image model is a machine learning model that may be used to receive a natural language description (e.g., text) as an input and generate an image that matches the description. Some text-to-image models may be used to generate collages of images by arranging existing component images from e.g., a database of clip art. Some text-to-image models may be able to generate more complex images such as compositions based on the input text (e.g., “an astronaut riding a horse”). Some techniques, such as style transfer, can combine two images (a first image and a second image) together so that the resultant output image retains core elements of the first image but appears to be painted in the style of the second image. Other techniques, such as prompt interpolation, includes interpolating two different text prompts in a text latent space (i.e., a representation of compressed data in which similar data points are closer together in space) before being used for image generation. When using prompt interpolation, in cases where the two concepts are extremely dissimilar (e.g., a living object and a non-living object), the generated image is typically dominated by one of the concepts.”).
and the generation code is configured to cause at least one of the at least one processor to select the predetermined number of image parts that include character strings that match or are similar to the plurality of attribute values included in the derived metadata, and generate the image by combining the predetermined number of image parts. (Liew [0031]: “As referenced herein, “similarity” may refer to a numeric value representing a degree of how close two objects (e.g., two images, two concepts corresponding to respective objects, etc.) are when the two objects are compared. It is to be understood that a similarity between two objects may be determined by using e.g., technologies such as sum of squared differences, mutual information, normalized mutual information, cross-correlation, etc. In an example embodiment, the higher the similarity (or value), the more contextually similar the two objects are. In such embodiment, a similarity (or value) “0” may indicate that the two objects are completely different. It is to be understood that in another example embodiment, the lower the similarity (or value), the more contextually similar the two objects are. In such embodiment, a similarity (or value) “0” may indicate that the two objects are identical.” Liew teaches mixing or blending two concepts, which is the predetermined number of image parts, to generate the image. Liew [0071]: “As shown in FIG. 4C, a de-nosing process may be used in the content generation process to generate a content of the output object (and/or to generate the output object itself including both the layout and the content), given a content conditioner (e.g., a conditional text prompt “coffee machine”, etc.) to mix or blend the two concepts (i.e., the source image at x.sub.0 and the object corresponding to the content conditioner). The de-noising process starts with a process node “{circumflex over (x)}.sub.K”. In an example embodiment, the process node {circumflex over (x)}.sub.k receives the layout image at x.sub.k as an input. From {circumflex over (x)}.sub.k, the de-noising process may de-noise (e.g., remove or filter noise such as Gaussian noise from) the layout image given the content conditioner to generate a de-noised, mixed, and/or blended image (e.g., at {circumflex over (x)}.sub.K-1). The de-noising process may de-noise the layout image (and mix or blend with the object corresponding to the content conditioner) from a previous process node and at {circumflex over (x)}.sub.0, the de-noised, mixed, and/or blended image becomes the output image. That is, the de-noising process may repeatedly de-noise the layout image at x.sub.k to generate a de-noised, mixed, and/or blended image after K operations or iterations ({circumflex over (x)}.sub.K.fwdarw.{circumflex over (x)}.sub.0).”).
As per claim 4, Liew anticipates the claimed:
4. The image generation apparatus according to claim 2, the program code further comprising: change code configured to cause at least one of the at least one processor to change at least a portion of the image, wherein each of the plurality of image parts stored in the storage is composed of a plurality of sub-parts, and each of the plurality of sub-parts is provided with an identifier for identifying the sub-part, and the change code is configured to cause at least one of the at least one processor to change a color of a sub-part corresponding to an identifier designated by a user. (Liew [0061]-[0062]: “[0061] At the optional block 440 (Apply weight), the processor may apply weighted text-image or image-text cross-attention (e.g., by changing a weight value s) when running the model (e.g., the diffusion process and/or the de-noising process), e.g., to the conditioner. As referenced herein, “attention” may refer to a technique that mimics cognitive attention. The effect of attention may enhance some parts of the input data while diminishing other parts. Learning which part of the data is to be enhanced rather than or before another depends on the semantic context. It is to be understood that text-image or image-text cross-attention may enable attention with context from both image and text, and may infer the text-image or image-text similarity by aligning image region and text features.
[0062] It is also to be understood that the processor may re-weigh (e.g., by changing a weight value s) the text-image or image-text cross-attention to increase or reduce the magnitude of a concept. For example, for a text-image cross-attention map M (a real data distribution in the space of the number of spatial and text tokens) and a conditional prompt y=“a photo of tiger”, the map M may be scaled corresponding to the “tiger” token with a value “s” while keeping the remaining map M unchanged. In an example embodiment, “s” may range from −D to D, with D being a number (for example, D=2, etc.). In an example embodiment, as shown in FIG. 5B (where s changes from 0.5 to 1.0 to 2.0), increasing “s” may result in more elements of “tiger” being injected into the synthesis (the generated output object).”
The text token are the image subparts that are portions of the image.).
As per claim 5, Liew anticipates the claimed:
5. The image generation apparatus according to wherein, when the metadata is modified by a user, the generation code causes at least one of the at least one processor to generate an image based on the modified metadata. (Liew [0001]: “The embodiments described herein pertain generally to generating an object using a diffusion model. More specifically, the embodiments described herein pertain to generating an object from text and/or image input using a pre-trained text-to-image diffusion-based generative model.”).
As per claim 13, Liew anticipates the claimed:
13. The image generation apparatus according to claim 1, wherein the derivation code causes at least one of the at least one processor to modify the derived metadata by changing an attribute value of the plurality of attribute values in response to a modification instruction input by the user, and the generation code causes at least one of the at least one processor to generate an updated image based on the modified metadata. (Liew [0054]: “In an example embodiment, to improve the mixing or blending quality, the processor may adjust (e.g., increase, decrease, etc.) the number k to control the content generation and/or provide flexibility of what the generated content may look like. For example, when the first input is an image of a rabbit (see the source image of FIG. 5A), and the second input is a “coffee machine” text prompt or text input, the mixing or blending ratio of the two concepts or objects can be controlled (e.g., more like a “rabbit” versus more like a “coffee machine”) by adjusting the number k of operations or iterations (or the point of time) at which or when the content generation process starts. If k is small (e.g., either as an absolute number or relative to T, e.g., when k is 5 while T is 100), most of the structural details of the rabbit may be preserved (since the number of operations or iterations for the diffusion process (e.g., k) is small) e.g., after executing the diffusing process for the number k operations or iterations, and there may be only a few operations or iterations left for the model to modify (e.g., for the de-noising process to modify the diffused rabbit image with the content conditioner “coffee machine”). If k is large (e.g., either as an absolute number or relative to T, e.g., when k is 85 while T is 100), most of the layout information of the rabbit may be corrupted e.g., after executing the diffusing process for the number k operations or iterations, and as a result, the model (e.g., the de-noising process of the model) may be given a higher flexibility to generate the desired content (e.g., modifying the diffused or corrupted rabbit image with the content conditioner “coffee machine”). In a case when k equals T, the semantic blending process may reduce to a text-to-image generation from a random Gaussian noise (e.g., the layout of the rabbit may completely disappear” Liew [0055]: “It is to be understood that a desired injection time (or a desired k, the desired number of operations or iterations) may be determined based on the determined semantic similarity value N between the two concepts. For example, as shown in the upper row of the images in FIG. 5A, when mixing or blending “rabbit” and “hamster” (two similar concepts or objects), the model may require fewer de-noising operations or iterations to generate features of the hamster due to the high level of similarity between a rabbit and a hamster (e.g., only need to modify the eyes and nose of the rabbit). As such, the number k may be adjusted (e.g., decreased) based on the determined similarity value N. In another example, as shown in the lower row of the images in FIG. 5A, when mixing or blending “rabbit” and “coffee machine” (two extremely dissimilar concepts or objects), the model may require more de-noising operations or iterations to generate the features of the coffee machine and/or to overwrite the rabbit details due to their high level of dissimilarity. As such, the number k may be adjusted (e.g., increased) based on the determined similarity value N.” The modification of the features of two dissimilar concepts is the changing of attribute values of the metadata.).
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claim 6 is rejected under 35 U.S.C. 103 as being unpatentable over Liew and further in view of Wang (Wang, Zijie J., et al. "Diffusiondb: A large-scale prompt gallery dataset for text-to-image generative models." Proceedings of the 61st annual meeting of the association for computational linguistics (volume 1: Long papers). 2023.).
As per claim 6, Liew alone does not explicitly teach the claimed limitations.
However, Liew in combination with Wang teaches the claimed:
6. The image generation apparatus according to claim 1, wherein the metadata is metadata written in a JSON (JavaScript Object Notation) format. (Liew teaches that the metadata is written in a format that lends itself to JSON format. It has objects expressed with keys or tags and associated values. Liew [0029]: “It is to be understood that an object (e.g., an image, etc.) may include metadata (such as keywords, tags, or descriptions associated with the object) and non-metadata such as features of the object (e.g., color, shape, texture, element or component or part, position of the element, or any other information that may be derived from the object itself) rather than the metadata. As referenced herein, the “layout” of an object may refer to the shape, the color, the texture, and/or the position of the element(s) of the object within the object. As referenced herein, the “content” of an object may refer to non-metadata information (e.g., the element, the component, the part, and/or the category of the object, etc.) of the object other than the layout.” Wang teaches a database of Object-Text pairs for image generation stored in a JSON format. Wang 2.4: “We organize DIFFUSIONDB using a flexible file structure. We first give each image a unique file name using Universally Unique Identifier (UUID, Version 4) (Leach et al., 2005). Then, we or ganize images into 14,000 sub-folders—each in cludes 1,000 images. Each sub-folder also includes a JSON file that contains 1,000 key-value pairs mapping an image name to its metadata. An exam ple of this image-prompt pair can be seen in Fig. 2. This modular file structure enables researchers to f lexibly use a subset of DIFFUSIONDB.”).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to use the JSON format for mapping images to prompts as taught by Wang with the system of Liew in order to store and transmit this information in a popular, standardized format.
Claim 7 is rejected under 35 U.S.C. 103 as being unpatentable over Liew in view of Saharia (US 20230377226 A1).
As per claim 7, Liew alone does not explicitly teach the claimed limitations.
However, Liew in combination with Saharia teaches the claimed:
7. The image generation apparatus according to claim 1, wherein the language model is an LLM (Large Language Model). (Saharia [0052]: “This specification introduces an image generation system that combines the power of text encoder neural networks (e.g., large language models (LLMs)) with a sequence of generative neural networks (e.g., diffusion-based models) to deliver text-to-image generation with a high degree of photorealism, fidelity, and deep language understanding. In contrast to prior work that use primarily image-text data for model training, a contribution described in this specification is that contextual embeddings from text encoders, pre-trained on text-only corpora, are effective for text-to-image generation.”).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to use the LLM for text-to-image generation as taught by Saharia with the system of Liew in order to use that method and structure of data processing to generate accurate images from text.
Claims 10-11 are rejected under 35 U.S.C. 103 as being unpatentable over Liew (Pub No. in view of Azarian (Pub No. US 20250166236 A1).
As per claim 10, Liew alone does not explicitly teach the claimed limitations.
However, Liew in combination with Azarian teaches the claimed:
10. The image generation apparatus according to claim 1, wherein the prompt includes an instruction to select a value that semantically matches the text information, from a plurality of selectable attribute values, for each of a plurality of keys. (Azarian [0065]: “FIG. 7 illustrates an example computation flow for generating cross-attention features between two data sets, such as sequences 702 and 704, as an example, using the guidance modules 330.sub.j-330.sub.j+n, and/or the guidance modules 402.sub.j-402.sub.j+n. The first sequence 702 represents the text embedding 316 for example. The second sequence 704 represents a patch 408 for example. Cross-attention is applied to the sequences 702, 704 to indicate how strongly associated each spatial region of the image (e.g., patch 408) is with different words in the text prompt (e.g., a specific token in the text embedding 316). As shown, value weights 706 are applied to the first sequence 702 to transform its features to value sequence 718. Key weights 708 are applied to the first sequence 702 to transform its features to key sequence 712. Query weights 710 are applied to the second sequence 704 to obtain query sequence 714. The key sequence 712 and query sequence 714 are compared, for instance using matrix multiplication, to generate an attention matrix 716. In certain aspects, this provides attention scores representing the relevance between specific portions of the two sequences 702, 704. The attention matrix 716 is applied to the value sequence 718 to generate cross-attended feature sequence 720 that represent an aggregation of relevant features from the first sequence 702 based on the second sequence 704. In certain aspects, the attention matrix 716 (or scores) may be modified as described above with respect to FIGS. 3-5.”).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to use the metadata attributes represented as keys as taught by Azarian with the system of Liew in order to organize the features of the image for better generation and accessibility. The motivation of claim 1 is incorporated herein.
As per claim 11, Liew alone does not explicitly teach the claimed limitations.
However, Liew in combination with Azarian teaches the claimed:
11. The image generation apparatus according to claim 1, wherein the language model is a GPT (Generative Pre-trained Transformer) model or a BERT (Bidirectional Encoder Representations from Transformers) model. (Azarian [0036]: “In certain aspects, the example image generation system 200, which may be similar to or the same as image generation system 100, generates an output image 202 based on one or more conditionings 204. The one or more conditionings 204 may include a text prompt 206 that is the same as or similar to the text prompt 104 of FIG. 1. The text prompt 206 can be provided to a text encoder 208 configured to encode the text prompt 206 into a text embedding 210. The text encoder 208 may be the same as or similar to the text encoder 110 of FIG. 1. The text encoder 208 encodes a text prompt 206 into a latent feature representation that captures the semantic meaning of the text. In certain aspects, the text encoder 208 uses a transformer model such as BERT, CLIP, or another language model, to encode the input text prompt 206 into an embedding vector. In certain aspects, the transformer encodes the text into an embedding vector (e.g., text embedding 210) by passing it through multiple self-attention layers. Each self-attention layer transforms the text into a progressively more abstract high-dimensional representation that extracts contextual relationships between the words and concepts in the text prompt 206. The resulting output of the text encoder 208 is a dense latent vector (e.g., text embedding 210) representation of the encoded text prompt 206. This text embedding 210 captures the semantic essence of the text prompt 206 in a format consumable by the image generation system 200.”).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to use the BERT model as taught by Azarian with the system of Liew in order to use that type of generative model to output images.
Claim 12 is rejected under 35 U.S.C. 103 as being unpatentable over Liew in view of Coyne (Pub No. US 7016828 B1).
As per claim 12, Liew alone does not explicitly teach the claimed limitations.
However, Liew in combination with Coyne teaches the claimed:
12. The image generation apparatus according to claim 1, wherein the derivation code is configured to cause at least one of the at least one processor to derive a plurality of sets of metadata by changing a parameter of the language model. (Coyne teaches explicitly changing the parameters of the language model used to make the images. Coyne col. 12 line 63-col. 13 line 20: “In the first example below, the depiction rule will be considered if either of the actions "kick" or "punt" are depicted. Furthermore, this particular depiction rule is an example of a depiction rule that might be used when there is no path or specified trajectory. An example of a sentence that indicates no path or specified trajectory might be John kicked the ball, as opposed to John kicked the ball over the fence. This exemplary depiction rule also checks to see that there is a direct object (in this case "ball") and that the size of the direct object is larger than four feet. If the object is smaller than four feet, then a second, possibly less restrictive, depiction rule may be used. Of course, the parameters evaluated by the depiction rule may be changed without departing from the scope of the invention. (define-depiction (:action ("kick" "punt") "in place") :test (and (not opath) direct-object (>(find-size-of direct-object) 4.0)) :fobjects (list (make-pose-depictor "kick" :actor subject) (make-spatial-relation-depictor "behind
" subject direct-object))) (define-depiction (:action ("kick" "punt")) :fobjects (make-path-verb-depictor subject t 3.0 "kick ball" direct-object opath :airborne-figure-p t))”).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to use the parameter change of the language model as taught by Coyne with the system of Liew in order to control the way the model evaluates the images to generate new images. The motivation of claim 1 is incorporated herein.
Conclusion
Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a).
A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to THOMAS JOHN FOSTER whose telephone number is (571)272-5053. The examiner can normally be reached Mon, Fri 8:30-6. Tues-Thurs 7:30-5.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Daniel Hajnik can be reached at 571-272-7642. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/THOMAS JOHN FOSTER/Examiner, Art Unit 2616
/HAI TAO SUN/Primary Examiner, Art Unit 2616