Prosecution Insights
Last updated: August 17, 2026
Application No. 18/344,481

IMPROVING IMAGE QUALITY VIA DISCRETE NATURAL LANGUAGE TOKENS

Non-Final OA §103§112
Filed
Jun 29, 2023
Examiner
DORVIL, RICHEMOND
Art Unit
2658
Tech Center
2600 — Communications
Assignee
International Business Machines Corporation
OA Round
1 (Non-Final)
33%
Grant Probability
At Risk
1-2
OA Rounds
4m
Est. Remaining
58%
With Interview

Examiner Intelligence

Grants only 33% of cases
33%
Career Allowance Rate
19 granted / 58 resolved
-29.2% vs TC avg
Strong +25% interview lift
Without
With
+25.2%
Interview Lift
resolved cases with interview
Typical timeline
3y 6m
Avg Prosecution
28 currently pending
Career history
91
Total Applications
across all art units

Statute-Specific Performance

§101
12.3%
-27.7% vs TC avg
§103
55.2%
+15.2% vs TC avg
§102
11.3%
-28.7% vs TC avg
§112
16.6%
-23.4% vs TC avg
Black line = Tech Center average estimate • Based on career data from 58 resolved cases

Office Action

§103 §112
DETAILED ACTION The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Specification The title of the invention is not descriptive. A new title is required that is clearly indicative of the invention to which the claims are directed. The following title is suggested: Sonar Image Generation Via Discrete Natural Language Tokens. The disclosure is objected to because of the following informalities: In ¶[0019], “by learn mappings” should be “by learning mappings”. In ¶[0057], Figure 8 illustrates “voice-to-text model 804”, but there is no reference numeral 804 described in the Specification. A voice-to-text model should have a reference numeral 804 in the Specification. In ¶[0058], “image 812” should be “image 814” as illustrated in Figure 5. (four occurrences) In ¶[0059], “image 812” should be “image 814” as illustrated in Figure 5. Appropriate correction is required. Claim Rejections - 35 USC § 112 The following is a quotation of 35 U.S.C. 112(b): (b) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention. The following is a quotation of 35 U.S.C. 112 (pre-AIA ), second paragraph: The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the applicant regards as his invention. Claims 1 to 20 are rejected under 35 U.S.C. 112(b) or 35 U.S.C. 112 (pre-AIA ), second paragraph, as being indefinite for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor (or for applications subject to pre-AIA 35 U.S.C. 112, the applicant), regards as the invention. Independent claims 1, 8, and 15 set forth limitations of an image generation component that generates an image “based on discrete natural language tokens” and “wherein the discrete tokens are not semantic”, which are indefinite under 35 U.S.C. §112(b). Generally, it appears to be inconsistent to claim that natural language is not semantic. That is, ‘semantics’ is defined as relating to meaning in language, and ‘natural language’ has meaning, so it does not appear to make sense to claim natural language as having tokens that are not semantic. Applicants’ invention as described, e.g., at ¶[0046]: Figure 5 of the Specification, takes sounds from sonar of ‘may oh oh feed heavy tuing plus’ and somehow converts these sounds into images. These sounds do not appear to be ‘natural language’ which is defined as language as it is spoken and written. Nor do all of these sounds necessarily represents words, e.g., ‘tuing’. Applicants may overcome this rejection by at least canceling the limitation of “wherein the discrete tokens are not semantic”. Here, any inconsistency between “natural language” and “not semantic” can be removed by deleting this limitation of “wherein the discrete tokens are not semantic”. However, Applicants should consider some alternate way of claiming this feature that does not involve natural language because these tokens as described do not necessarily correspond to natural language as it is conventionally understood by one skilled in the art. Claim Rejections - 35 USC § 103 The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claims 1 to 2, 7 to 9, 14 to 16, and 20 are rejected under 35 U.S.C. 103 as being unpatentable over Patterson et al. (U.S. Patent Publication 2005/0270905) in view of Diamos et al. (U.S. Patent Publication 2018/0061439). Concerning independent claims 1, 8, and 15, Patterson et al. discloses a system including a processor and a memory, a method, and a computer program product, comprising: “an image generating component that uses a first neural network model to generate an image of an environment detected by an imaging sonar, based on [discrete tokens in natural language that represent] sound waves reflected by structures in the environment[, wherein the tokens are not semantic]” – sidescan sonar is an acoustic imaging technology that uses high frequency sounds waves to illuminate the sea floor and produce realistic pictures; as sound waves propagate away from the sidescan transducers, objects in the path of the beam reflect some of the acoustic energy back to the transducer and these signals are passed on to a video display (¶[0004]); a sonar image of sonar targets within a liquid medium are collected and identified; image transformation is performed on the image using an extraction algorithm to generate a feature vector related to sonar targets, and the generated feature vector is presented to a neural network to classify the image with respect to the sonar targets (“an image generation component that uses a first neural network to generate an image of the environment detected by an imaging sonar”) (¶[0008] - ¶[0010]); as sound waves propagate away from sidescan transducers 1, objects in the path of the beam reflect some of the acoustic energy back into the sonar instrument (“based on . . . sound waves reflected by the structures in the environment”); sidescan sonar can capture images including a shark 32, a tire 23, or a downed aircraft 24 within a liquid medium (¶[0033] - ¶[0034]: Figure 3); particle analysis is performed on extracted regions of interest (ROIs) to obtain a feature vector, and feature vectors related to sonar targets are presented to a neural network to classify the image (¶[0042] - ¶[0043]: Figure 2: Step 240 to 250); once neural network 700 is trained with prototypes or ground-truth images, it is ready to perform recognition tasks on previously unseen data; classification consists of evaluating whether an N-dimensional input vector lies within an active influence field (AIF) of any prototype of the network, and if it is, then the input is recognized as belonging to that corresponding category (¶[0060]: Figure 11). Concerning independent claims 1, 8, and 15, Patterson et al. discloses all of these limitations with the exception of generating an image based on “discrete tokens in natural language” and “wherein the tokens are not semantic”. Arguably, Patterson et al.’s feature vectors can be construed as ‘discrete tokens’, and feature vectors are “not semantic” because they do not have meaning in natural language. However, Patterson et al. does not disclose discrete tokens in “natural language”. Here, Patterson et al. processes a sonar image with a neural network to label objects within the image, e.g., identifies a species of fish in a sonar image of fish. Concerning independent claims 1, 8, and 15, Diamos et al. teaches automatic audio captioning of a non-speech sound using a neural network. A non-speech sound is received, and relevant features are extracted from a raw audio waveform using a recurrent neural network (RNN) acoustic model, and a discrete sequence of characters represented in a natural language (“based on discrete tokens in natural language that represent sound waves”) is generated based on the relevant features, wherein the discrete sequence of characters comprises a caption that describes the non-speech sound. (Abstract; ¶[0004]) A caption generated for a raw audio signal including a dog barking may be ‘a dog barks four times’. The captions may describe speech events by translating speech into text. A caption generated for a raw audio may be ‘a dog barks four times’. (¶[0014]) Relevant features are extracted from a raw audio waveform using a recurrent neural network, and a discrete sequence of characters in natural language is generated based on the relevant features, where the discrete sequence of characters comprises a caption that describes the generated non-speech sound. (¶[0017]: Figure 1) One embodiment relates to captioning sounds of ocean waves crashing on rocks. (¶[0026]) Here, identifying feature vectors as a non-speech sound of a dog barking provides “wherein the discrete tokens are not semantic.” That is, non-speech sounds of a dog barking do not represent semantic speech with meaning. An objective is to provide automatic audio captioning of non-speech sounds. (¶[0002] - ¶[0004]) It would have been obvious to one having ordinary skill in the art to generate a sonar image with a neural network in Patterson et al. based on discrete tokens in natural language that are not semantic as taught by Diamos et al. for a purpose of providing automatic audio captioning of non-speech sounds. Concerning claims 2, 9, and 16, Patterson et al. discloses “a first neural network” for generating an image with identified objects, and Diamos et al. teaches “a second neural network” for converting sound waves to discrete tokens in natural language, e.g., to convert sound waves of a barking dog into a natural language description of a barking dog. Concerning claims 7, 14, and 20, Diamos et al. teaches comparing a generated label to a reference label using one or more cost functions that indicates an accuracy of the neural network. A number and size of each neural network layer is chosen until an RNN acoustic model 160 and/or an RNN language model 180 achieves a desired level of accuracy (“with quality above a second defined threshold”). (¶[0027]: Figure 1B) Claims 3 to 5, 10 to 12, and 17 to 18 are rejected under 35 U.S.C. 103 as being unpatentable over Patterson et al. (U.S. Patent Publication 2005/0270905) in view of Diamos et al. (U.S. Patent Publication 2018/0061439) as applied to claims 1, 8, and 15 above, and further in view of Neumann et al. (U.S. Patent Publication 2022/0171043). Concerning claims 3, 10, and 17, Patterson et al. does not disclose the limitations of “a data enhancement component that adds one or more new tokens to a token sequence comprising the discrete tokens, wherein the one or more new tokens respectively represent one or more features in the environment not captured by the imaging sonar.” Concerning claims 3, 10, and 17, Neumann et al. teaches a sonar display that receives sonar data and additional data from a data source other than a sonar transducer and determines object characteristics using sonar data and additional data to determine an estimated object type for an object represented with the sonar data with a sonar image generated based on the sonar data. (Abstract) Improved display features may be provided with artificial intelligence techniques so that object characteristics and/or estimated object types may be determined and presented in a clear manner. The operations include receiving additional data from a data source other than the one or more sonar transducer assemblies including a geographic area, a time of day, or time of year. A model is developed using machine learning and artificial intelligence based on historical comparisons of historical sonar data and historical additional data. (¶0007] - ¶[0010]) Additional data includes geographical data from maps or nautical charts, temperature data, and time data. Additional data may be provided from a variety of sources including from a camera, a radar, a thermometer, a clock, a pressure sensor, a direction sensor, or a position sensor. (¶[0054]) Additional data may be used alongside available sonar data to develop and improve a model that may predict an outline of objects and distinguish between two different types of objects. (¶[0067]) Figures 3B to 3E present embodiments of display 300 that includes first window 308 comprising an estimated object-type, e.g., ‘salmon fish’. (¶[0080] - ¶[0083]: Figure 3B to 3E) Neumann et al., then, teaches “a data enhancement component that adds one or more new tokens to a token sequence” that “represent one or more features in the environment not captured by the imaging sonar.” These features of additional data are “new tokens” that are added to data presented to a machine learning or artificial intelligence algorithm corresponding to a neural network of Patterson et al. An objective is to present sonar data using artificial intelligence so that all users including novices can interpret sonar data without a significant amount of knowledge and experience. (¶[0002] - ¶[0003]) It would have been obvious to one having ordinary skill in the art to add additional data representing features in an environment not captured by an imaging sonar as taught by Neumann et al. as tokens input to a neural network of Patterson et al. for a purpose of presenting sonar data so that all users including novices can interpret sonar data without a significant amount of knowledge and experience. Concerning claims 4 to 5, 11 to 12, and 18, Neumann et al. teaches that additional data may be used alongside available sonar data to develop and improve a model that may predict an outline of objects and distinguish between two different types of objects (“wherein addition of the one or more new tokens to the token sequence enables the one or more features to be included in the image of the environment”). (¶[0067]) Broadly, additional data that is used alongside sonar data is “token injection enhancement.” That is object identification in sonar images is ‘enhanced’ by adding (’injecting’) additional data from non-sonar sources that are features presented as ‘tokens’ and a ‘token sequence’ to a neural network of a machine learning algorithm. Claims 6, 13, and 19 are rejected under 35 U.S.C. 103 as being unpatentable over Patterson et al. (U.S. Patent Publication 2005/0270905) in view of Diamos et al. (U.S. Patent Publication 2018/0061439) as applied to claims 1, 8, and 15 above, and further in view of Thorsley et al. (U.S. Patent Publication 2022/0374766). Broadly, Patterson et al. can be construed to disclose feature vectors (“the discrete tokens”) of sound waves that “act as a bridge between the sound waves and the image of the environment” to classify an image with respect to targets of interest. That is, feature vectors are “discrete tokens . . that represent sound waves”, and these feature vectors “act as a bridge” between the sound waves and the image. Similarly, Diamos et al. can be construed to teach the limitation of “wherein the discrete tokens act as a bridge between the sound waves and the image of the environment” because sound waves of non-speech sounds of a barking of a dog are converting into a discrete sequence of natural language characters, e.g., ‘a dog barks four times’, that provides a caption representing “the image of the environment.” Diamos et al. does not expressly teach the limitation of “and cause a computational load on at least the first neural network model to fall below a first defined threshold.” However, Diamos et al. teaches that a decoder algorithm includes a beam search component and includes a maximum beam size parameter which determines the maximum number of labels to search simultaneously. (¶[0027]: Figure 1B) The number of layers, size of filters in a given layer, and a number of filters in a given layer are hyper-parameters that may be tuned, and a performance improvement should be balanced against computational limits, i.e., increasing the layer count, filter count, and/or filter size may result in an audio captioning system that requires too much time to train. (¶[0035]: Figure 2A) Generation of concept vectors for each sample in an audio frame can lead to excessive computation performed over audio frames that correspond to silence, but attention addresses the problems related to per-sample generation of concept vectors. (¶[0044] - ¶[0045]: Figure 3A) Diamos et al., then, teaches that “a computational load” can increase with a complexity of a machine learning model, but that this might be addressed by an attention mechanism. Thorsley et al. teaches learned threshold token pruning for transformer neural networks which prunes tokens from an input sequence below a predetermined threshold score. (Abstract) A self-attention mechanism scales quadratically with a length of an input sequence so that longer sequences require more computation. Computation involved during deep neural network training is reduced by reducing redundant tokens as data passes through different blocks of a transformer deep-learning model by using threshold-based pruning. (¶[0025]) Here, reducing computation by pruning until a desired threshold score is obtained is “cause a computational load on at least the first neural network to fall below a first defined threshold.” An objective is to reduce an amount of computation involved during deep neural network training. (¶[0038]: Figure 2) It would have been obvious to one having ordinary skill in the art to cause a computational load on a neural network to fall below a defined threshold as taught by Thorsley et al. due to the complexity of a machine learning model of Diamos et al. for a purpose of reducing an amount of computation involved during deep neural network training. Conclusion The prior art made of record and not relied upon is considered pertinent to Applicants’ disclosure. Haley et al., Bengio et al., Mei et al., Clark et al., Ji et al., Laster, Husain, Radkoff et al., and Donderici disclose related prior art. Any inquiry concerning this communication or earlier communications from the examiner should be directed to MARTIN LERNER whose telephone number is (571) 272-7608. The examiner can normally be reached Monday-Thursday 8:30 AM-6:00 PM. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Richemond Dorvil can be reached at (571) 272-7602. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /MARTIN LERNER/Primary Examiner Art Unit 2658 May 2, 2025
Read full office action

Prosecution Timeline

Jun 29, 2023
Application Filed
Nov 20, 2023
Response after Non-Final Action
Jul 15, 2026
Non-Final Rejection mailed — §103, §112 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12688362
TRANSLATION DEVICE
3y 2m to grant Granted Jul 21, 2026
Patent 12658188
SYSTEM AND METHOD FOR ROBOT INITIATED PERSONALISED CONVERSATION WITH A USER
2y 8m to grant Granted Jun 16, 2026
Patent 12651593
INTENT RECOGNITION METHODS, APPARATUSES, AND DEVICES
3y 1m to grant Granted Jun 09, 2026
Patent 12591738
Autocorrect Candidate Selection
3y 1m to grant Granted Mar 31, 2026
Patent 12573397
ELECTRONIC APPARATUS AND CONTROLLING METHOD THEREOF
2y 6m to grant Granted Mar 10, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
33%
Grant Probability
58%
With Interview (+25.2%)
3y 6m (~4m remaining)
Median Time to Grant
Low
PTA Risk
Based on 58 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month