Prosecution Insights
Last updated: October 02, 2026
Application No. 18/588,278

INFORMATION PROCESSING DEVICE, INFORMATION PROCESSING METHOD, AND COMPUTER PROGRAM PRODUCT

Final Rejection §103
Filed
Feb 27, 2024
Priority
Jul 04, 2023 — JP 2023-110203
Examiner
KOETH, MICHELLE M
Art Unit
2671
Tech Center
2600 — Communications
Assignee
Kabushiki Kaisha Toshiba
OA Round
2 (Final)
77%
Grant Probability
Favorable
3-4
OA Rounds
0m
Est. Remaining
93%
With Interview

Examiner Intelligence

Grants 77% — above average
77%
Career Allowance Rate
339 granted / 442 resolved
+14.7% vs TC avg
Strong +16% interview lift
Without
With
+16.5%
Interview Lift
resolved cases with interview
Typical timeline
2y 2m
Avg Prosecution
33 currently pending
Career history
473
Total Applications
across all art units

Statute-Specific Performance

§101
6.1%
-33.9% vs TC avg
§103
69.5%
+29.5% vs TC avg
§102
7.8%
-32.2% vs TC avg
§112
10.3%
-29.7% vs TC avg
Black line = Tech Center average estimate • Based on career data from 442 resolved cases

Office Action

§103
DETAILED ACTION Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Response to Arguments Applicant’s arguments and amendments in the Amendment filed August 11, 2026 (herein “Amendment”) with respect to the objection to the Title have been fully considered and are persuasive. The objection to the Title has been withdrawn. Applicant’s arguments and amendments in the Amendment, with respect to the objections to claims 3, 6 and 8, and therefore any claims depending therefrom, have been fully considered and are persuasive. The objections to claims 3, 6 and 8, and therefore any claims depending therefrom, has been withdrawn. Applicant’s amendments in the Amendment with respect to the invocation of interpretation under 35 U.S.C. 112(f) have been fully considered and are persuasive. Accordingly, claim 1 and various claims depending therefrom are no longer invoking interpretation under 35 U.S.C. 112(f). Applicant’s arguments and amendments in the Amendment, with respect to the rejection of claims 3 and 9 under 35 U.S.C. 112(b) for indefiniteness have been fully considered and are persuasive. The rejection of claims 3 and 9 under 35 U.S.C. 112(b) has been withdrawn. Applicant’s arguments and amendments in the Amendment, with respect to the rejection of claims 1–10 under 35 U.S.C. 101 for being directed to an abstract idea without a practical application or significantly more have been fully considered and are persuasive. The rejection of claims 1–10 under 35 U.S.C. 101 has been withdrawn. Applicant’s arguments and amendments in the Amendment, with respect to the rejection of claims 1–10 under 35 U.S.C. 103 have been fully considered and are persuasive. Therefore, the rejection has been withdrawn. However, upon further consideration, a new ground of rejection is made in view of Kim et al., US Patent Application Publication No. US 2015/0019532 A1 and Chen et al., US Patent Application Publication No. US 2017/0124432 A1. Claim Rejections - 35 USC § 103 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows: 1. Determining the scope and contents of the prior art. 2. Ascertaining the differences between the prior art and the claims at issue. 3. Resolving the level of ordinary skill in the pertinent art. 4. Considering objective evidence present in the application indicating obviousness or nonobviousness. Claims 1, 6–7, 9 and 10 are rejected under 35 U.S.C. 103 as being unpatentable over Costabello et al., US Patent No. US 10,949,718 B2 (herein “Costabello”) in view of Kim et al., US Patent Application Publication No. US 2015/0019532 A1 (herein “Kim”) in view of Chen et al., US Patent Application Publication No. US 2017/0124432 A1 (herein “Chen”). Regarding claims 1, 9 and 10, with deficiencies of Costabello noted in square brackets [], and significant differences between the claims noted with curly brackets {}, Costabello teaches {an information processing device comprising: one or more hardware processors configured to – claim 1 / A computer program product comprising a non-transitory computer-readable medium including programmed instructions, the instructions causing a computer to perform a method comprising: - claim 9 / An information processing method, implemented by an information processing device, comprising: - claim 10} (Costabello col. 15, ll. 5–25, system 100 including hardware and software combinations, the hardware including a processor, where col. 3, ll. 4–6 and 21–26 teaches a system processing an input image and query and outputting a response (information processing)): detect a plurality of pieces of different object information that each include an object area containing an object to be detected (Costabello col. 3, ll. 31–34, and col. 3, l. 47–col. 4, l. 2, computer vision techniques are applied to generate symbolic representations of information in the image such as the pixel locations (plural) in the image that identify a region (object area) of the input image that corresponds to an object shown in the image, such as a plural objects cat, dog, and trees) and object identification information for identifying the object to be detected, from an image (Costabello col. 3, ll. 53–64, symbolic representations including content classification for detected content in the image which can include type or category of an object); generate [a plurality of object images, by cutting out the object areas corresponding to the detected pieces of different object information] from the image (Costabello col. 4, ll. 10–23, col. 6, ll. 15–37, and col. 12, ll. 7–28, image-processing framework, further described in fig. 2, encodes a portion of the image as a sub-symbolic feature vector/embedding, the portion being for example that of a cat (object) in the image, therefore “cutting” out of the image the cat object by generating the sub-symbolic embedding data for the pixels including the cat region of the image); [acquire at least one question specific to the object identification information for the detected pieces of different object information]; and perform a visual question answering (VQA) process with the at least one question, for each of the plurality of object images (Costabello fig. 1, col. 5, l. 38–col. 6, l. 14, at inference time, an inference controller receives an inference query (question) and provides a natural language answer regarding the inquired object, for example an input query of “Which animal in this image is able to climb trees,” and the answer being provided as “The cat can climb the trees.”) [by converting each object image into a feature vector by an image encoder, converting each acquired question into a feature vector by a text encoder, inputting each converted object image and converted acquired question into an artificial intelligence (AI) model and obtaining an answer from the AI model]. While Costabello teaches separately encoding a cat portion of an image, thereby teaching at least generating one object image by cutting out an area from the image, Costabello does not teach, where Kim teaches generate a plurality of object images, by cutting out the object areas corresponding to the detected pieces of different object information from the image (Kim ¶¶61, 65–66, 71, an image is divided into multiple two-dimensional areas later stored as individual images, based on colors, shapes, textures and contours corresponding to identified objects displayed in respective areas). Further, Costabello does not, but Kim teaches acquire at least one question specific to the object identification information for the detected pieces of different object information (Kim fig. 7, ¶¶74–75, a query (question) is received (acquire) with words corresponding to (specific to)keywords of objects (object identification information) in an image that found from previously being stored on the server (the detected pieces of different object information)). Still further, Costabello does not where Chen teaches by converting each object image into a feature vector by an image encoder (Chen fig. 4, ¶65, ABC-CNN architecture (image encoder) extracts an image feature map from an input image), converting each acquired question into a feature vector by a text encoder (Chen fig. 4, ¶65, obtaining a dense question embedding (feature vector) from an input question using a long short term memory layer (text encoder)), inputting each converted object image and converted acquired question into an artificial intelligence (AI) model and obtaining an answer from the AI model (Chen fig. 4, ¶65, an answer to the question is generated in step 430 based on a fusion of the image feature map, and the deep question embedding by using a multi-class classifier with attention weights (AI model)). Therefore, taking the teachings of Costabello and Kim together as a whole, it would have been obvious to a person having ordinary skill in the art before the effective filing date of the claimed invention to have modified the image processing of Costabello to include the image dividing based on object characteristics as disclosed in Kim at least because doing so would provide for a more accurate image search result. See Kim ¶¶2, 44. Further, taking the teachings of Costabello as modified and Chen together as a whole, it would have been obvious to a person having ordinary skill in the art before the effective filing date of the claimed invention to have modified the image processing of Costabello to include the question acquisition and image feature and question feature processing as disclosed in Chen at least because doing so would improve the performance of a visual question-answering system. See Chen ¶29. Regarding claim 6, Costabello teaches wherein the one or more hardware processors are configured to assign question identification information to at least one question applied to each object image (Costabello col. 5, ll. 19–62, fig. 1, inference query assigned parameters of ?, has_skill, climb_trees to an input unstructured query that corresponds to an input image with objects), and output VQA process result information in which the object identification information, the question identification information, and an answer to a question identified by the question identification information are associated with one another (Costabello col. 5, l. 57–col. 6, l. 14, inference query parameters are input to the inference controller to generate an embedding query with the content classifications (associating object identification information with the question identification information) and identifying a specific multi-modal embedding that is a best replacement for the ? parameter of the inference query parameters, thus associating a multi-modal embedding that defines the embedding result as an inference response (answer to a question) with the embedding query and content classifications). Regarding claim 7, Costabello teaches wherein the one or more hardware processors are configured to display, on a display device, display information in which the VQA process result information is assigned to an object to be detected identified by the object identification information included in the VQA process result information (Costabello col. 14, ll. 1–6, the system displays the natural language response on a display through a graphical user interface, where col. 6, l. 5–14, teaches that the natural language response is the result of the conversion of the inference response generated as a result of processing the image and input query (question) regarding (assigned to) a detected object in the image (such as “cat” having a skill of tree climbing)). Claims 2–4, are rejected under 35 U.S.C. 103 as being unpatentable over Costabello in view of Kim, in view of Chen, as set forth above regarding claim 1, further in view of Pham et al., US Patent Application Publication No. US 2022/0129693 A1 (herein “Pham”). Regarding claim 2, with deficiencies of Costabello noted in square brackets [], Costabello teaches further comprising a storage unit configured to store [at least one question] (Costabello col. 4, ll. 46–62, fig. 1, multi-modal embedding model shown in fig. 1 as a database symbol, where col. 8, ll. 64–66 teach the multi-modal embeddings being stored in the multi-modal embedding model, and where col. 15, ll. 41–49 teaches that all of the system and its logic and data structures (of which the multi-modal embedding model would be understood to be a data structure) is stored on non-transitory computer readable storage media) for each object type indicating a type of object to be detected, wherein the object identification information includes information indicating an object type (Costabello col. 4, ll. 48–50, multi-modal embeddings include an aggregation of the symbolic embeddings and the sub-symbolic embeddings, where col. 3, ll. 50–60 teach the symbolic embeddings including a type of object detected in the content of the image), and the one or more hardware processors are configured to acquire [at least one question] according to the object type included in the object identification information, [from the storage unit] (Costabello col. 5, ll. 9–18, multi-modal embedding framework identifies (acquires) an embeddings result set having specific multi-modal embeddings based on an input query from the multi-modal embedding model). Costabello does not but Pham teaches storage including at least one question from the storage unit (Pham ¶¶39–40 and 42, fig. 3, question and answer acquisition unit extracts a set of questions (acquire at least one question) from a table storing the questions (storage) based on an image feature, where ¶99 teaches the image feature calculated from a region of the image corresponding to an object). Therefore, taking the teachings of Costabello as modified above and Pham together as a whole, it would have been obvious to a person having ordinary skill in the art before the effective filing date of the claimed invention to have modified the image processing of Costabello to include the question acquisition unit acquiring questions and question storage as disclosed in Pham at least because doing so would allow for answers to be provided that are sourced in official authoritative guides/manuals such as a work manual or safety manual, and thus provide more accurate answers. See Pham ¶52. Regarding claim 3, Costabello teaches wherein the one or more hardware processors are configured to detect (Costabello col. 3, ll. 31–34, and col. 3, l. 47–col. 4, l. 2, computer vision techniques are applied to generate symbolic representations of information in the image such as the pixel locations in the image that identify a region of the input image that corresponds to an object shown in the image) the object identification information, from the image, the object area containing an object to be detected of an object type read out from the storage unit (Costabello col. 5, ll. 9–18, col. 4, ll. 48–50, multi-modal embeddings include an aggregation of the symbolic embeddings and the sub-symbolic embeddings, where the multi-modal embedding framework identifies (acquires) an embeddings result set having specific multi-modal embeddings based on an input query from the multi-modal embedding model (storage unit), and where noted above the symbolic representations include regions (area) corresponding to detected object, and object types). Regarding claim 4, Costabello teaches wherein the one or more hardware processors are configured to receive a question sentence on the image from a user, identify the object type from the question sentence, and detect at least one piece of object information including an object area and the object identification information, the object area containing an object to be detected of the identified object type (Costabello fig. 1, col. 5, ll. 9–18, and col. 5, l. 57–col. 6, l .5, the multi-modal embedding framework 116 identifies an embeddings result set having specific multi-modal embeddings based on an input query received (question sentence on the image from the user) from the multi-modal embedding model (storage unit), and where noted above the symbolic representations include regions (area) corresponding to detected object, and object types). Claim 5 is rejected under 35 U.S.C. 103 as being unpatentable over Costabello in view of Kim in view of Chen view of Pham, as set forth regarding claim 2, further in view of Kodama, US Patent No. US 12,361,670 B2 (herein “Kodama”). Regarding claim 5, Costabello as modified above does not teach, but Kodama teaches wherein the one or more hardware processors are configured to transform the object area according to at least one of the object type and the question sentence (Kodama col. 10, ll. 3–12, numbers 200–209 as functional blocks realized by processors (unit), where first variable magnification unit enlarges or reduces (transform) image data of a target region (object area), where col. 10, ll. 39–42 teaches that the target region is determined through object detection before the first magnification, the object detection also obtaining the category of the object (according to object type)), and generate at least one object image, by cutting out at least one object area or at least one object area transformed by the one or more hardware processors, from the image (Kodama col. 10, ll. 13–15, fig. 8, the image cutting-out unit cuts out the region of the target object from the region division map on which semantic segmentation is performed, after the first variable magnification unit enlarges or reduces). Therefore, taking the teachings of Costabello as modified above and Kodama together as a whole, it would have been obvious to a person having ordinary skill in the art before the effective filing date of the claimed invention to have modified the image processing of Costabello to include the enlarging/reducing and the cutting-out unit as disclosed in Kodama at least because doing so would allow for appropriately setting the size of processing stages to accommodate the size of the image, thereby reducing the need for increased calculation and memory size. See Kodama col. 10, ll. 55–63. Claim 8 is rejected under 35 U.S.C. 103 as being unpatentable over Costabello in view of Kim in view of Chen, further in view of Bharadwaj et al., US Patent Application Publication No. US 2024/0242029 A1 (herein “Bharadwaj”). Regarding claim 8, with deficiencies of Costabello noted in square brackets [], Costabello teaches wherein the image is a frame included in a moving image (Costabello col. 3, ll. 11–12, the input image as a video frame), and the one or more hardware processors are configured to detect at least one piece of object information from the frame (Costabello col. 6, ll. 15–37, image features are extracted from the image and content classifications for the detected objects are associated to pixel regions), [and vote an answer to a question obtained by performing the VQA process on the object image generated for each frame, and determine an answer to the question for each object image, based on a result of vote]. Costabello as modified above does not teach, where Bharadwaj teaches and vote an answer to a question obtained by performing the VQA process on the object image generated for each frame, and determine an answer to the question for each object image, based on a result of the vote (Bharadwaj ¶¶160–166, VQA process resulting in answers which are scored by a common sense scorer (CSS) module 602 that also applies a majority voting approach, where the majority vote wins to determine the best answer to a question). Therefore, taking the teachings of Costabello as modified above and Bharadwaj together as a whole, it would have been obvious to a person having ordinary skill in the art before the effective filing date of the claimed invention to have modified the image processing of Costabello to include the voting on answers as disclosed in Bharadwaj at least because doing so would allow for more accurate and stable predictions in a VQA answer. See Bharadwaj ¶166 and Abstract. Conclusion Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a). A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action. Any inquiry concerning this communication or earlier communications from the examiner should be directed to MICHELLE M KOETH whose telephone number is (571)272-5908. The examiner can normally be reached Monday-Thursday, 09:00-17:00, Friday 09:00-13:00, EDT/EST. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Vincent Rudolph can be reached at 571-272-8243. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. MICHELLE M. KOETH Primary Examiner Art Unit 2671 /MICHELLE M KOETH/Primary Examiner, Art Unit 2671
Read full office action

Prosecution Timeline

Feb 27, 2024
Application Filed
Apr 13, 2026
Non-Final Rejection mailed — §103
Jul 16, 2026
Examiner Interview Summary
Jul 16, 2026
Applicant Interview (Telephonic)
Aug 11, 2026
Response Filed
Sep 16, 2026
Final Rejection mailed — §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12749311
DISPLAY APPARATUS THAT PROVIDES ANSWER TO QUESTION BASED ON IMAGE AND CONTROLLING METHOD THEREOF
2y 11m to grant Granted Sep 29, 2026
Patent 12743872
MULTI-OBJECTIVE DENSE OPEN-VOCABULARY IMAGE RECORDING
2y 5m to grant Granted Sep 22, 2026
Patent 12731591
APPARATUS AND METHOD FOR MDCT M/S STEREO WITH GLOBAL ILD WITH IMPROVED MID/SIDE DECISION
2y 10m to grant Granted Sep 08, 2026
Patent 12705775
METHOD AND APPARATUS FOR OBTAINING 3D INFORMATION OF VEHICLE
3y 11m to grant Granted Aug 11, 2026
Patent 12700397
SOUND OUTPUT CONTROL DEVICE, SOUND OUTPUT CONTROL METHOD, AND SOUND OUTPUT CONTROL PROGRAM
2y 11m to grant Granted Aug 04, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

3-4
Expected OA Rounds
77%
Grant Probability
93%
With Interview (+16.5%)
2y 2m (~0m remaining)
Median Time to Grant
Moderate
PTA Risk
Based on 442 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month