DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Priority
Foreign priority is not claimed.
Information Disclosure Statement
The information disclosure statement (IDS) submitted on 01/22/2025 is in compliance with the provisions of 37 CFR 1.97. Accordingly, the information disclosure statement is being considered by the examiner.
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102 of this title, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claims 1-4, 9-10, 11-14 and 19 are rejected under 35 U.S.C. 103 as being unpatentable over Kharbanda et al. (US Patent No. 11978271 B1) in view of Ramesh et al. (US Patent No. 11983806 B1).
Regarding Claim 1,
Kharbanda discloses A method comprising: receiving a command for a vision-language model to extract target information from an initial image data structure, wherein the vision-language model is incapable of extracting the target information within a predetermined degree of accuracy; (Kharbanda, Col. 5, Lines 52-65, Col., 6. Lines 1-15, discloses Vision language models can leverage learned image and language associations to generate natural language captions for images; however, vision language models can struggle with details including object particularity. The lack of particularity can lead to the generation of generalized queries and/or prompts, which may fail to provide results that are specific to and/or applicable to the features depicted in the image. For example, a user may provide an image with a question “how do I take care of this?” The vision language model may process the image to determine the image depicts a plant, which can be leveraged to generate a refined query of “what do plants need to stay alive and grow?” The refined query can be processed to determine search results that may be associated with general care instructions for plants, which may include watering twice a week, half a day of direct sunlight, and loamy soil. However, the generalized care instructions may not be suitable for the specific plant depicted in the image (e.g., a succulent (e.g., an agave plant) needs less water and different soil, and a shuttlecock fern may thrive in shade over direct sunlight). Therefore, the utilization of generalized information for the object class may be detrimental to the caretaking and counter to the original purpose of the inputs; a command for vision language is obtained and processed to determine which specific visual details to be extracted from an input image and sometimes may not result in output of similar exact query when in visual language processing)
executing a super-resolution machine learning model on the initial image data structure to output an enhanced image data structure; (Kharbanda, Col. 6, Lines 16-25, Col. 6, Lines 35-50, Col. 6, Lines 60-67, Col. 7, Lines 1-15, discloses Pairing instance level object recognition with vision language model processing can be utilized to generate detailed captions, queries, and/or prompts. Combining scene understanding with instance understanding can be leveraged for image searching, image indexing, automated content generation and/or understanding, and/or other image understanding tasks. For example, the augmented language output can be leveraged as and/or to generate a detailed query and/or a detailed prompt to obtain and/or generate additional information. The particularity can lead to improved tailoring of search results and/or generative prompts; Multimodal large language models (e.g., large vision language models) may be tuned and/or trained to have a rough understanding of images. For example, processing an image with a large language model may be able to output “This is a black dog sitting on a beach”. However, object recognition systems can be trained and/or configured to recognize objects at instance level granularity. In the same image, the object recognition system can recognize the dog's breed as Australian Kelpie and the beach as Bondi beach. When coupling the two systems, the systems and methods can teach and/or condition the large language model that the scene includes an Australian Kelpie sitting on the Bondi beach. The large language model can then learn and/or be prompted to describe a scene at instance level granularity. The systems and methods disclosed herein can be utilized to recognize every product in an aisle as the user walks past the products and can then help the user to find the products that meet their dietary restrictions and/or other preferences and criteria; systems and methods disclosed herein can be leveraged to process a plurality of different data types (e.g., image data, text data, video data, audio data, statistical data, graph data, latent encoding data, and/or multimodal data) to generate outputs that may be in a plurality of different data formats (e.g., image data, text data, video data, audio data, statistical data, graph data, latent encoding data, and/or multimodal data). For example, the input data may include a video that can be processed to generate a summary of the video, which may include a natural language summary, a timeline, a flowchart, an audio file in the form of a podcast, and/or a comic book. The object recognition system can be leveraged for object specific details, while the scene understanding model (e.g., a vision language model) may be leveraged for scene recognition and/or frame group understanding. In some implementations, one or more additional models may be leveraged for context understanding. For example, a hierarchical video encoder may be utilized for frame understanding, frame sequence understanding, and/or full video understanding. Audio input processing may include the utilization of a text-to-speech model, which may be implemented as part of the language model; advanced and trained models trained on granular details (super resolution models) are used extract precise object (target) within a scene (image data structure) and context-based prompt is generated based on understanding of the scene and target object detail data)
generating a prompt, wherein the prompt includes the enhanced image data structure and an instruction to extract the target information from the enhanced image data structure; (Kharbanda, Col. 7, Lines 1-15, discloses systems and methods of the present disclosure provide a number of technical effects and benefits. As one example, the system and methods can be utilized to generate instance level scene recognition outputs. In particular, the systems and methods disclosed herein can leverage a vision language model in parallel with an object recognition system to generate a natural language output that is both scene-aware and object specific. The augmented language output can then be utilized as a query for a search and/or a prompt for generative model content generation; enhanced prompt is generated based on detail context trained model and specific object (target) is extracted from a scene image (image data structure))
executing the vision-language model on the prompt to output the target information with at least the predetermined degree of accuracy; (Kharbanda, Col. 7, lines 48-61, Fig. 1, discloses an example detailed image captioning system 10 according to example embodiments of the present disclosure. In some implementations, the detailed image captioning system 10 is configured to receive, and/or obtain, a set of input data including image data 12 descriptive of an environment with one or more objects and, as a result of receipt of the image data 12, generate, determine, and/or provide an augmented language output 22 that is descriptive of an object instance level scene recognition. Thus, in some implementations, the detailed image captioning system 10 can include a vision language model 18 that is operable to perform scene recognition and an object recognition block 14 that is operable to perform object recognition; (26) In particular, the detailed image captioning system 10 can obtain input data, which can include image data 12 descriptive of one or more input images. The one or more input images can be descriptive of an environment and one or more objects. The environment can include a room, a landscape, a city, a town, a sky, and/or other environments. In some implementations, the environment is descriptive of a user environment generated with one or more image sensors of a user computing device. The one or more objects can include products, people, plants, animals, art pieces, structures, landmarks, and/or other objects; vision language model outputs target description as requested by prompt) and
presenting the target information. (Kharbanda, Col. 20, Lines 39-48, discloses at 810, the computing system can determine one or more search results associated with the augmented language output. The one or more search results can be associated with one or more web resources. In some implementations, determining the one or more search results associated with the augmented language output can include determining a plurality of search results are responsive to a search query including the augmented language output. Additionally and/or alternatively, the computing system can provide the plurality of search results for display; target description result is output on display screen)
Kharbanda does not explicitly disclose wherein the enhanced image data structure comprises a higher pixel resolution than the initial image data structure;
Ramesh discloses wherein the enhanced image data structure comprises a higher pixel resolution than the initial image data structure; (Ramesh, Col. 6, Lines 51-65, Col. 7, Lines 36-45, discloses Disclosed embodiments may improve the technical field of AI-based image generation and editing, including achieving images that are high-resolution, new compositions, and/or photorealistic. Generated images may maintain the semantics, context, meaning, and style of an original image while editing one or more portions of the image. Disclosed embodiments may also enable generation of an image beyond the image's original borders, thereby providing larger images which can capture more details. Disclosed embodiments may provide the capability for a user to guide the model, such as allowing the user to give input to the model and thereby exert more control over the AI-based generated images, resulting in an image that more closely resembles the image desired by the user; Illustrative embodiments of the present disclosure are described below. In one embodiment a system may include at least one memory storing instructions and at least one processor configured to execute the instructions to perform operations for regenerating a region of an image with a machine learning model based on a text input. Generating an image may include at least one of producing, creating, making, computing, calculating, deriving, or outputting digital information (e.g., pixel information, such as one or more pixel values), which may form an image. Regenerating an image may refer to generating an image, such as making edits, alterations, or additions to an image. As referenced herein, an image may include digital images, pictures, art, and/or digital artwork. Text or text inputs may include written language, natural language, printed language, description, captions, prompts, sequence of characters, sentences, and/or words. In some embodiments, a text description may include a text input (e.g., received from an input/output device); the newly generated images with use of improved prompt are higher resolution with detailed description of scene or image)
Therefore, it would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention was made to combine the teachings of Kharbanda in view of Ramesh having a method of creating a prompt that better describes a target object in an image and outputs enhanced target image, with the teachings of Ramesh having, by the system of describing a target object in an image with detailed contextual description with higher resolution with a prompt in order to produce high quality image in applications including art fields.
Regarding Claim 2,
The combination of Kharbanda and Ramesh further wherein the initial image data structure defines a single image and the enhanced image data structure defines a single enhanced image, and wherein the single enhanced image is formed from the single image. (Kharbanda, Col. 1. Lines 39-61, discloses one example aspect of the present disclosure is directed to a computer-implemented method. The method can include obtaining, by a computing system including one or more processors, image data. The image data can include an input image. The method can include processing, by the computing system, the input image with an object recognition model to generate a fine-grained object recognition output. The fine-grained object recognition output can be descriptive of identification details for an object depicted in the input image. The method can include processing, by the computing system, the input image with a vision language model to generate a language output. The language output can include a set of predicted words predicted to be descriptive of the input image. In some implementations, the set of predicted words can include a coarse-grained term descriptive of predicted identification of the object depicted in the input image. The method can include processing, by the computing system, the fine-grained object recognition output and the language output to generate an augmented language output. The augmented language output can include the set of predicted words with the coarse-grained term replaced with the fine-grained object recognition output; single image is input and enhanced image is output as one image data structure being processed and output as enhanced image).
Regarding Claim 3,
The combination of Kharbanda and Ramesh further discloses wherein the super-resolution machine learning model comprises at least one of a single-scale architecture or a multi-scale architecture, wherein the single-scale architecture processes the initial image data structure at a single super-resolution scale and the multi-scale architecture processes the initial image data structure at a plurality of scales. (Ramesh, Col. 11, Lines 53-67, Col. 12, Lines 1-5, Col. 12, Lines 39-67, Col. 13, Lines 1-13, discloses some embodiments of a machine learning model may involve one or more sub-models. Some disclosed embodiments involve a first sub-model configured to generate an image embedding based on at least one of the text input, the masked image, and the masked region. An embedding, including an image embedding, may include an output of a machine learning model, such as a numeric, vector, or spatial representation of the input to the text encoder. For example, an image embedding may include a mapping of a given image input to a multidimensional vector representation, such as a lower-dimensional representation of the image. In some embodiments, the first sub-model may include a prior model alongside an image encoder that is jointly trained with the prior model. The image encoder may receive inputs of the masked image and the masked region, and the prior model may receive an input of text, such as a caption, or a text embedding. The first sub-model may generate an image embedding, such as an image embedding corresponding to the masked image. Some disclosed embodiments involve a second sub-model configured to generate the enhanced image based on at least one of the image embedding, the text input, the masked image, and the masked region. The second sub-model may include a decoder and/or a diffusion model, as described herein. For example, the second sub-model may be provided the image embedding and may generate an output, such as the enhanced image, based on the image embedding. It will be appreciated that using image embeddings for the second sub-model to create the enhanced image by replacing pixel values in the masked region provides various advantages to the generated image, including producing more realistic images. In some examples, the image embedding and the text input, such as a caption, may be provided to the second sub-model. It will be appreciated that providing the caption of the desired image enhancement prompt to the second sub-model may provide additional guidance to the second sub-model for image generation, such as in the diffusion process, thereby increasing the accuracy of the generated enhanced image. For example, a given enhancement caption may assist the model and enable the model to generate pixel values for the image segment which more accurately correspond to the caption; the machine learning model may include a deep learning model alongside a large-language model. Large-language models, which may include one or more transformers, may refer to models capable of utilizing large datasets to understand, generate, and predict content. For example, transformers may be deep learning architectures, such as neural networks, which include encoder and/or decoder networks, as well as attention mechanisms, to learn contextual relationships within data. In some embodiments, the machine learning model has a text-to-image model trained on a set of images. A text-to-image model may refer to any machine learning model configured to generate images with guidance of a text input. For example, text-to-image machine learning models may generate an output image corresponding to a text input, such as a caption or image generation prompt, as well as generate modified versions of an image based on text inputs, such as text prompts to alter or edit features within the input; architecture model includes multiple artificial intelligence layers of input and output to generate and execute prompt outputs). Additionally, the rational and motivation to combine the references Kharbanda and Ramesh as applied in rejection of claim 1 apply to this claim.
Regarding Claim 4,
The combination of Kharbanda and Ramesh further discloses wherein the super-resolution machine learning model is at least one of an enhanced deep super-resolution (EDSR) network, a multi-scale deep super-resolution system, and a super-resolution residual network. (Ramesh, Col. 11, Lines 53-67, Col. 12, Lines 1-5, Col. 12, Lines 39-67, Col. 13, Lines 1-13, some embodiments of a machine learning model may involve one or more sub-models. Some disclosed embodiments involve a first sub-model configured to generate an image embedding based on at least one of the text input, the masked image, and the masked region. An embedding, including an image embedding, may include an output of a machine learning model, such as a numeric, vector, or spatial representation of the input to the text encoder. For example, an image embedding may include a mapping of a given image input to a multidimensional vector representation, such as a lower-dimensional representation of the image. In some embodiments, the first sub-model may include a prior model alongside an image encoder that is jointly trained with the prior model. The image encoder may receive inputs of the masked image and the masked region, and the prior model may receive an input of text, such as a caption, or a text embedding. The first sub-model may generate an image embedding, such as an image embedding corresponding to the masked image. Some disclosed embodiments involve a second sub-model configured to generate the enhanced image based on at least one of the image embedding, the text input, the masked image, and the masked region. The second sub-model may include a decoder and/or a diffusion model, as described herein. For example, the second sub-model may be provided the image embedding and may generate an output, such as the enhanced image, based on the image embedding. It will be appreciated that using image embeddings for the second sub-model to create the enhanced image by replacing pixel values in the masked region provides various advantages to the generated image, including producing more realistic images. In some examples, the image embedding and the text input, such as a caption, may be provided to the second sub-model. It will be appreciated that providing the caption of the desired image enhancement prompt to the second sub-model may provide additional guidance to the second sub-model for image generation, such as in the diffusion process, thereby increasing the accuracy of the generated enhanced image. For example, a given enhancement caption may assist the model and enable the model to generate pixel values for the image segment which more accurately correspond to the caption; the machine learning model may include a deep learning model alongside a large-language model. Large-language models, which may include one or more transformers, may refer to models capable of utilizing large datasets to understand, generate, and predict content. For example, transformers may be deep learning architectures, such as neural networks, which include encoder and/or decoder networks, as well as attention mechanisms, to learn contextual relationships within data. In some embodiments, the machine learning model has a text-to-image model trained on a set of images. A text-to-image model may refer to any machine learning model configured to generate images with guidance of a text input. For example, text-to-image machine learning models may generate an output image corresponding to a text input, such as a caption or image generation prompt, as well as generate modified versions of an image based on text inputs, such as text prompts to alter or edit features within the input; multiple of deep neural network systems are implemented as deep layer artificial intelligence network model). Additionally, the rational and motivation to combine the references Kharbanda and Ramesh as applied in rejection of claim 1 apply to this claim.
Regarding Claim 9,
The combination of Kharbanda and Ramesh further discloses after executing the vision-language model, receiving an additional command for at least one of the vision-language model and a language model to format the target information into a structured language data structure. (Kharbanda, Col. 5, lines 52-67, Col. 6, Lines 1-15, discloses systems and methods disclosed herein can process an image with a vision language model and a fine-grained object recognition model in parallel to generate an output that is scene-aware and object-aware while being formatted in a natural language format. The parallel processing can be separate and independent such that the scene-aware output and the object-aware output are determined separately and without influence of the other. Token replacement can be utilized to replace coarse-grained object recognition (e.g., object class recognition (e.g., a plant, a human, a car, a building, etc.)) of the vision language model with the fine-grained recognition (e.g., specific object recognition indicating the particular object identification (e.g., a Tiger lily, George Washington, a Model T soft-top convertible with 5L engine, Monticello, etc.)) of the instance level object recognition system. For example, the systems and methods can include processing the input image with an object recognition system to generate an object recognition output descriptive of identification details for the particular object depicted in the input image. The identification details can include an instance-level identification descriptive a specific and detailed identification for the object. The systems and methods can also process the input image with the vision language model to generate a language output descriptive of scene recognition for the entire scene depicted in the input image. The scene recognition may be less particular than the object recognition output. Therefore, the systems and methods may process the object recognition output and the language output to generate an augmented language output that leverages the scene recognition of the language output and the particularity of the object recognition output; object (target) description is formatted into a natural vision language prompt for detailed description for enhanced target data structure to be output). Additionally, the rational and motivation to combine the references Kharbanda and Ramesh as applied in rejection of claim 1 apply to this claim.
Regarding Claim 10,
The combination of Kharbanda and Ramesh further discloses receiving a second command for the vision-language model to determine that the a resolution of the initial image data structure, wherein the vision-language model is incapable of extracting the target information with a predetermined degree of accuracy at the determined resolution; and generating, in response to determining the resolution of the initial image data structure, a third command for the vision-language model to execute the super-resolution machine learning model on the initial image data structure. (Kharbanda, Col. 6, Lines 16-25, Lines 35-53, Lines 60-67, Col. 7, Lines 1-15, discloses Pairing instance level object recognition with vision language model processing can be utilized to generate detailed captions, queries, and/or prompts. Combining scene understanding with instance understanding can be leveraged for image searching, image indexing, automated content generation and/or understanding, and/or other image understanding tasks. For example, the augmented language output can be leveraged as and/or to generate a detailed query and/or a detailed prompt to obtain and/or generate additional information. The particularity can lead to improved tailoring of search results and/or generative prompts; Multimodal large language models (e.g., large vision language models) may be tuned and/or trained to have a rough understanding of images. For example, processing an image with a large language model may be able to output “This is a black dog sitting on a beach”. However, object recognition systems can be trained and/or configured to recognize objects at instance level granularity. In the same image, the object recognition system can recognize the dog's breed as Australian Kelpie and the beach as Bondi beach. When coupling the two systems, the systems and methods can teach and/or condition the large language model that the scene includes an Australian Kelpie sitting on the Bondi beach. The large language model can then learn and/or be prompted to describe a scene at instance level granularity. The systems and methods disclosed herein can be utilized to recognize every product in an aisle as the user walks past the products and can then help the user to find the products that meet their dietary restrictions and/or other preferences and criteria; systems and methods disclosed herein can be leveraged to process a plurality of different data types (e.g., image data, text data, video data, audio data, statistical data, graph data, latent encoding data, and/or multimodal data) to generate outputs that may be in a plurality of different data formats (e.g., image data, text data, video data, audio data, statistical data, graph data, latent encoding data, and/or multimodal data). For example, the input data may include a video that can be processed to generate a summary of the video, which may include a natural language summary, a timeline, a flowchart, an audio file in the form of a podcast, and/or a comic book. The object recognition system can be leveraged for object specific details, while the scene understanding model (e.g., a vision language model) may be leveraged for scene recognition and/or frame group understanding. In some implementations, one or more additional models may be leveraged for context understanding. For example, a hierarchical video encoder may be utilized for frame understanding, frame sequence understanding, and/or full video understanding. Audio input processing may include the utilization of a text-to-speech model, which may be implemented as part of the language model; advanced and trained models trained on granular details (super resolution models) are used extract precise object (target) within a scene (image data structure) and context-based prompt is generated based on understanding of the scene and target object detail data).
Claims 11-14 and 19 recite system with elements corresponding to the method steps recited in Claim 1. Therefore, the recited elements of the system claims 11-14 and 19 are mapped to the proposed combination in the same manner as the corresponding steps of Claims 1-4 and 9. Additionally, the rationale and motivation to combine the Kharbanda and Ramesh references presented in rejection of Claim 1, apply to these claims.
Furthermore, the combination of Kharbanda and Ramesh further discloses A system comprising: a computer processor; a data repository in communication with the computer processor, wherein the data repository stores: a command, target information, an initial image data structure, a predetermined degree of accuracy, an enhanced image data structure, and a prompt comprising the enhanced image data structure and an instruction (Kharabanda, Fig. 9A, Col. 24, Lines 3-55, discloses user computing system 102 and/or the server computing system 130 can train the models 120 and/or 140 via interaction with the third party computing system 150 that is communicatively coupled over the network 180. The third party computing system 150 can be separate from the server computing system 130 or can be a portion of the server computing system 130. Alternatively and/or additionally, the third party computing system 150 may be associated with one or more web resources, one or more web platforms, one or more other users, and/or one or more contexts; third party computing system 150 can include one or more processors 152 and a memory 154. The one or more processors 152 can be any suitable processing device (e.g., a processor core, a microprocessor, an ASIC, a FPGA, a controller, a microcontroller, etc.) and can be one processor or a plurality of processors that are operatively connected. The memory 154 can include one or more non-transitory computer-readable storage mediums, such as RAM, ROM, EEPROM, EPROM, flash memory devices, magnetic disks, etc., and combinations thereof. The memory 154 can store data 156 and instructions 158 which are executed by the processor 152 to cause the third party computing system 150 to perform operations. In some implementations, the third party computing system 150 includes or is otherwise implemented by one or more server computing devices; network 180 can be any type of communications network, such as a local area network (e.g., intranet), wide area network (e.g., Internet), or some combination thereof and can include any number of wired or wireless links. In general, communication over the network 180 can be carried via any type of wired and/or wireless connection, using a wide variety of communication protocols (e.g., TCP/IP, HTTP, SMTP, FTP), encodings or formats (e.g., HTML, XML), and/or protection schemes (e.g., VPN, secure HTTP, SSL); machine-learned models described in this specification may be used in a variety of tasks, applications, and/or use cases; the input to the machine-learned model(s) of the present disclosure can be image data. The machine-learned model(s) can process the image data to generate an output. As an example, the machine-learned model(s) can process the image data to generate an image recognition output (e.g., a recognition of the image data, a latent embedding of the image data, an encoded representation of the image data, a hash of the image data, etc.). As another example, the machine-learned model(s) can process the image data to generate an image segmentation output. As another example, the machine-learned model(s) can process the image data to generate an image classification output. As another example, the machine-learned model(s) can process the image data to generate an image data modification output (e.g., an alteration of the image data, etc.). As another example, the machine-learned model(s) can process the image data to generate an encoded image data output (e.g., an encoded and/or compressed representation of the image data, etc.). As another example, the machine-learned model(s) can process the image data to generate an upscaled image data output. As another example, the machine-learned model(s) can process the image data to generate a prediction output).
Allowable Subject Matter
Claims 5-8 and 15-18 are objected to as being dependent upon a rejected base claim, but would be allowable if rewritten in independent form including all of the limitations of the base claim and any intervening claims.
Claims 20 is allowed over prior art.
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure:
US-20250259340-A1 (Cheng et al., The method (600) involves obtaining (605) text
prompt describing an element and an attribute value for a continuous attribute of the element. The text prompt is embedded (610) to obtain a text embedding in a text embedding space. An attribute value is embedded (615) to obtain an attribute embedding in the text embedding space using a continuous control model. A synthetic image is generated (620) using an image generation model based on the text embedding and the attribute embedding, where the synthetic image depicts continuous attribute of the element based on the attribute value and the continuous attribute comprises a three-dimensional characteristic of the element, Abstract)
US-20250329079-A1 (Zhou et al., The method (600) involves obtaining an input image
and a text prompt comprising an image modification request (605). A text response is generated (610) based on the input image and the text prompt using a language generation model, where the text response describes a modification to the input image corresponding to an input modification request. A synthetic image is generated (615) based on the input image and an output embedding of the language generation model using an image generation model, where the synthetic image depicts the modification to the input image, Abstract)
Any inquiry concerning this communication or earlier communications from the examiner should be directed to PINALBEN V PATEL whose telephone number is (571)270-5872. The examiner can normally be reached M-F: 10am - 8pm.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Chineyere Wills-Burns can be reached at 571-272-9752. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/Pinalben Patel/Examiner, Art Unit 2673