DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Status of the Claims
Claims 1, 3-11, 13, 15-18 and 20 have been amended.
Claims 2, 12, 14 and 19 have been previously presented.
Claim Rejections - 35 USC § 112
The following is a quotation of the first paragraph of 35 U.S.C. 112(a):
(a) IN GENERAL.—The specification shall contain a written description of the invention, and of the manner and process of making and using it, in such full, clear, concise, and exact terms as to enable any person skilled in the art to which it pertains, or with which it is most nearly connected, to make and use the same, and shall set forth the best mode contemplated by the inventor or joint inventor of carrying out the invention.
Claims 1-20 are rejected under 35 U.S.C. 112(a) or 35 U.S.C. 112 (pre-AIA ), first paragraph, as failing to comply with the written description requirement. The claim(s) contains subject matter which was not described in the specification in such a way as to reasonably convey to one skilled in the relevant art that the inventor or a joint inventor, or for applications subject to pre-AIA 35 U.S.C. 112, the inventor(s), at the time the application was filed, had possession of the claimed invention. The following subject matter in claims 1, 11 and 20 was not described in the applicant’s originally filed Specification: “…determining a scene composition of the scene comprising a spatial arrangement of the plurality of image tiles…blend the plurality of image stiles based upon the scene prompt and the scene composition to generate the scene.“. In particular the ‘spatial arrangement’ in claims 1, 11 and 20, as well as to blend the image tiles based upon the ‘scene composition’ in claims 1, 11 and 20, are not mentioned in the Specification.
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102 of this title, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
This application currently names joint inventors. In considering patentability of the claims the examiner presumes that the subject matter of the various claims was commonly owned as of the effective filing date of the claimed invention(s) absent any evidence to the contrary. Applicant is advised of the obligation under 37 CFR 1.56 to point out the inventor and effective filing dates of each claim that was not commonly owned as of the effective filing date of the later invention in order for the examiner to consider the applicability of 35 U.S.C. 102(b)(2(c) for any potential 35 U.S.C. 102(a)(2) prior art against the later invention.
The factual inquiries set forth in Graham v. John Deere Co., 383 U.S. 1, 148 USPQ 459 (1966), that are applied for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows:
Determining the scope and contents of the prior art.
Ascertaining the differences between the prior art and the claims at issue.
Resolving the level of ordinary skill in the pertinent art.
Considering objective evidence present in the application indicating obviousness or nonobviousness.
Claims 1-3, 11-13 and 20 are rejected under 35 U.S.C. 103 as being unpatentable over Neilson et al. (“Solving jigsaw puzzles using image features”, Pattern Recognition Letters 29 (2008), p.1924-1933) in view of Paumard et al. (“Deepzzle: Solving Visual Jigsaw Puzzles With Deep Learning and Shortest Path Optimization”, IEEE Transactions on Image Processing, Vol. 29, 2020) and further in view of Sadr et al. (11,941,678).
Regarding claim 1, Neilson teaches a computer-implemented method for generating a scene (e.g., a method for automatic solving of the jigsaw puzzle problem based on using image features instead of the shape of the pieces. Neilson: Abstract L.1-2. The puzzle problem is the problem of assembling a jigsaw puzzle so that all pieces fit together forming a picture. Neilson: sec. 1.1 para. 1 L.1-2. The picture produces a scene, for example, Fig. 1(a) landscape; Fig. 1(b) construction and Fig. 10 Benjamin.),
PNG
media_image1.png
334
830
media_image1.png
Greyscale
PNG
media_image2.png
480
638
media_image2.png
Greyscale
the method comprising:
obtaining a plurality of image tiles based upon at least one user input (e.g., We assume that the pieces each have four sides and are arranged in a rectangular grid. Neilson: sec. 1.1 para. 1 L.2-4. Preprocessing generally consists of scanning real puzzles, separating the individual pieces and then rotating them to achieve the proper alignment. Neilson: sec. 1.3 para. 2 L.1-3. The puzzle pieces in grids are taken as tiles of the image, where each piece contributes portion of the whole picture of the puzzle);
detecting at least one other image tile included in the plurality of image tiles (e.g.,
PNG
media_image3.png
1226
616
media_image3.png
Greyscale
Neilson: sec. 3.1);
determining a scene composition of the scene comprising a spatial arrangement (sec. 1.4 1st para. line 1 – 2nd para. line 6) of the plurality of image tiles based upon the spatial positioning of each image tile included in the plurality of image tiles (e.g., The primary success criteria of the classification with a given set of classification features is to achieve a grouping of pieces, so that each group contains pieces which are likely to be inter-connected. Neilson: sec. 5.3 para. 1. In order to find the best feature sets for the two images landscape and construction, we had to try all combinations since we have no model for the correlations between features, but we decided to test color and texture features separately to reduce the total number of tests. Neilson: sec. 5.3 para. 2. The images indicates scenes of landscape and construction);
obtaining a scene prompt associated with the scene (e.g., Two complete sets of pieces (24 and 54 pieces) can be found at the (DIKU Image Group’s FTP server, 2008). Neilson: sec. 4 para. 3 L.6-8. The images landscape and construction provide two scenes. See 1_2 below); and
causing to blend the plurality of image tiles based upon the scene prompt and the scene composition to generate the scene (e.g., We have developed a new method for edge matching using only image features, making them independent of the border shape of the puzzle pieces. Our singlepiece algorithm is the first successful attempt at solving jigsaw puzzles without using shape information. On computer generated puzzles, we were able to solve puzzles of up to 320 pieces – larger than the record of 200 held by Goldberg et al. (2004), but with which it cannot be directly compared. In addition we managed to solve a real puzzle of 54 pieces, which is on par with other algorithms using image information. Neilson: sec. 6 para. 1. See 1_3 below).
While Neilson does not explicitly teach, Paumard teaches: (1_1). detecting a spatial positioning of each image tile included in the plurality of image tiles relative to at least one other image tile included in the plurality of image tiles (e.g., Figure 2 illustrates our puzzle-solving process. For each image in our dataset, we extract a square that we cut into 9 pieces. To mimic the erosion, we then randomly crop a fragment inside each piece, making sure there is a wide gap between the fragments. Then, we pair the central fragment with each lateral fragment. Each couple is processed by a neural network that predicts their relative position among the 8 alternatives. These probabilities are used to build a graph, in which we compute the shortest path to reassemble the puzzle. Paumard: sec. III A para. 1. We propose two extensions of this problem. First, we consider the case where the central fragment is unknown. In this case, we compute the relative positions supposing that each fragment is the central one. Then, we apply the shortest path algorithm in each of these graphs, and we select the most probable solution. Second, we deal with missing fragments and outsider fragments, which are frequent in archaeology. In this case, we allow fragments to be unused and positions to be unfilled. Paumard: sec. III A para. 2. The fragments of Paumard is interpreted to be the pieces of Neilson and the placement of pieces would include the positioning of the pieces of Paumard); and
(1_3). causing a machine learning model to blend the plurality of image tiles based upon the scene prompt to generate the scene (e.g., We tackle the image reassembly problem with wide space between the fragments, in such a way that the patterns and
colors continuity is mostly unusable. The spacing emulates the erosion of which the archaeological fragments suffer. We crop-square the fragments borders to compel our algorithm to learn from the content of the fragments. We also complicate the image
reassembly by removing fragments and adding pieces from other sources. We use a two-step method to obtain the reassemblies: 1) a neural network predicts the positions of the fragments despite the gaps between them; 2) a graph that leads to the best reassemblies is made from these predictions. In this paper, we notably investigate the effect of branch-cut in the graph of reassemblies. We also provide a comparison with the literature, solve complex images reassemblies, explore at length the dataset, and propose a new metric that suits its specificities. Paumard: Abstract. In [3], we proposed a preliminary method to tackle the puzzle-solving task with deep neural networks and graphs. We focused on solving 3 × 3 jigsaw puzzles made of same-sized squared 2D fragments (Figure 1), using a 2-step method. First, given a central fragment, we used a neural network to predict the relative position of each remaining fragment. Then, the best solution is obtained using a graph of the possible reassembly. Paumard: sec. 1 para. 2. Therefore, the automatic solving of the jigsaw puzzle problem of Neilson can be performed with a neural network (machine learning algorithm));
Therefore it would have been obvious to a person of ordinary skill in the art before the effective filing date of the claimed invention to combine the teaching of Paumard into the teaching of Neilson so that possible positions are considered to place new piece.
While the combined teaching of Neilson and Paumard does not explicitly teach, Sadr teaches: (1_2). obtaining a scene prompt associated with the scene (e.g., Systems and methods for searching using machine-learned model-generated outputs can provide a user with a medium for generating a theoretical dataset that can then be matched to a real world example. The systems and methods can include selecting a plurality of terms, which can be utilized to generate a prompt input that can be processed by a dataset generation model to generate a plurality of model-generated datasets. A selection can then be received that selects a particular model-generated database to utilize to query a database . Sadr: Abstract. Therefore, the prompt input is used to select the puzzle that is generated from the DIKU Image Group’s FTP server (2008));
Therefore it would have been obvious to a person of ordinary skill in the art before the effective filing date of the claimed invention to combine the teaching of Sadr into the combined teaching of Neilson and Paumard so that puzzles are loaded using input prompt to the DIKU Image Group’s FTP server.
Regarding claim 2, the combined teaching of Neilson, Paumard and Sadr teaches wherein the scene prompt comprises a textual prompt indicating how the plurality of image tiles should be blended to generate the scene (e.g., The method can include obtaining, by a computing system including one or more processors, a multi-modal prompt input. The multi-modal prompt input can include a prompt image and prompt text. The prompt image can be descriptive of a particular object. In some implementations, the prompt text can be descriptive of one or more particular details of the prompt image to augment. The method can include processing, by the computing system, the prompt image and the prompt text with an image generation model to generate a model-generated image. The model-generated image can be descriptive of a model-generated object. In some implementations, the model-generated object can be descriptive of the particular object augmented based on the prompt text. The method can include providing, by the computing system, the model-generated-image to a search engine as a search query based on one or more user inputs and receiving, by the computing system, one or more search results from the search engine based on the model-generated image. Sadr: c.3 L.19-37).
Regarding claim 3, the combined teaching of Neilson, Paumard and Sadr wherein obtaining the plurality of image tiles based upon the at least one user input comprises:
obtaining an image tile textual prompt that describes a desired characteristic of a corresponding image tile included in the plurality of image tiles (e.g., Current image generation systems utilize a prompt input box for receiving freeform text to be processed to generate one or more images. Sadr: c.1 L.42-44. The prompt image can be descriptive of a particular object. In some implementations, the prompt text can be descriptive of one or more particular details of the prompt image to augment. Sadr: c.3 L.24-27. Preprocessing generally consists of scanning real puzzles, separating the individual pieces and then rotating them to achieve the proper alignment. Neilson: sec. 1.3 para. 2 L.1-3);
providing the image tile textual prompt to the machine learning model (e.g., providing the prompt input 206 to a machine-learned model 208 (e.g., a dataset generation model) to receive a plurality of machine-learned model outputs (e.g., a plurality of model-generated datasets), Sadr: c.15 L.43-47); and
executing the machine learning model to generate the corresponding image tile based upon the textual prompt (e.g., providing the prompt input 206 to a machine-learned model 208 (e.g., a dataset generation model) to receive a plurality of machine-learned model outputs (e.g., a plurality of model-generated datasets), obtaining a selection input, providing the selected machine-learned model output to a search engine 216, and receiving one or more search results. Sadr: c.15 L.43-49).
Regarding claims 11-13, the claims are non-transitory computer-readable media claims of method claims 1-3 respectively. The claims are similar in scope to claims 1-3 respectively and they are rejected under similar rationale as claims 1-3 respectively.
Sadr teaches that “Systems and methods for searching using machine-learned model-generated outputs can provide a user with a medium for generating a theoretical dataset that can then be matched to a real world example. The systems and methods can include selecting a plurality of terms, which can be utilized to generate a prompt input that can be processed by a dataset generation model to generate a plurality of model-generated datasets. A selection can then be received that selects a particular model-generated database to utilize to query a database.” (Sadr: Abstract) and “ The system can include one or more processors and one or more non-transitory computer-readable media that collectively store instructions that, when executed by the one or more processors, cause the computing system to perform operations.” (Sadr: c.1 L.60-64).
Regarding claim 20, the claim is a system claim of method claim 1. The claim is similar in scope to claim 1 and it is rejected under similar rationale as claim 1. Sadr teaches that “Systems and methods for searching using machine-learned model-generated outputs can provide a user with a medium for generating a theoretical dataset that can then be matched to a real world example. The systems and methods can include selecting a plurality of terms, which can be utilized to generate a prompt input that can be processed by a dataset generation model to generate a plurality of model-generated datasets. A selection can then be received that selects a particular model-generated database to utilize to query a database.” (Sadr: Abstract).
Claims 4-5 and 14-15 are rejected under 35 U.S.C. 103 as being unpatentable over Neilson in view of Paumard and Sadr as applied to claim 1 (11) and further in view of Gafni et al. (“Make-A-Scene: Scene-Based Text-to-Image Generation with Human Priors”, arXiv:2203.13131, March 24, 2022, p.1-17; IDS).
Regarding claim 4, the combined teaching of Neilson, Paumard and Sadr teaches the computer-implemented method of claim 1, wherein obtaining the plurality of image tiles based upon the at least one user input comprises:
obtaining within a graphical user interface a sketch of a portion of the scene (see 4_1 below);
providing the sketch to the machine learning model (e.g., providing the prompt input 206 to a machine-learned model 208 (e.g., a dataset generation model) to receive a plurality of machine-learned model outputs (e.g., a plurality of model-generated datasets), Sadr: c.15 L.43-47); and
executing the machine learning model to generate a corresponding image tile based upon the sketch (e.g., providing the prompt input 206 to a machine-learned model 208 (e.g., a dataset generation model) to receive a plurality of machine-learned model outputs (e.g., a plurality of model-generated datasets), obtaining a selection input, providing the selected machine-learned model output to a search engine 216, and receiving one or more search results. Sadr: c.15 L.43-49).
While the combined teaching of Neilson, Paumard and Sadr does not explicitly teach, Gafni teaches: (4_1). obtaining within a graphical user interface a sketch of a portion of the scene (e.g., Methods that rely on text inputs only are more confined to generate within the training distribution, as demonstrated by [41]. Unusual objects and scenarios can be challenging to generate, as certain objects are strongly correlated with specific structures, such as cats with four legs, or cars with round wheels. The same is true for scenarios. “A mouse hunting a lion” is most likely not a scenario easily found within the dataset. By conditioning on scenes in the form of simple sketches, we are able to attend to these uncommon objects and scenarios, as demonstrated in Fig. 3, despite the fact that some objects do not exist as categories in our scene (mouse, lion). We solve the category gap by using categories that may be close in certain aspects (elephant instead of mouse, cat instead of lion). In practice, for non-existent categories, several categories could be used instead. Gafni: sec 4.7 para. 1 and Fig. 3; reproduced below for reference.
PNG
media_image4.png
484
816
media_image4.png
Greyscale
Therefore, a mouse is sketched to hunt a lion, a car is sketched with triangular wheels and continuous tracks are sketched on a bicycle);
Therefore it would have been obvious to a person of ordinary skill in the art before the effective filing date of the claimed invention to combine the teaching of Gafni into the combined teaching of Neilson, Paumard and Sadr so that uncommon objects and scenarios can be realized by sketching.
Regarding claim 5, the combined teaching of Neilson, Paumard and Sadr teaches the computer-implemented method of claim 1, further comprising:
obtaining within a graphical user interface a sketch of a portion of the scene (see 5_1 below);
obtaining an image tile textual prompt associated with the sketch of the portion of the scene, wherein the image tile textual prompt specifies a desired characteristic of the respective image tile (see 5_2 below); and
executing the machine learning model to generate the respective image tile based upon the sketch and the image tile textual prompt (e.g., providing the prompt input 206 to a machine-learned model 208 (e.g., a dataset generation model) to receive a plurality of machine-learned model outputs (e.g., a plurality of model-generated datasets), obtaining a selection input, providing the selected machine-learned model output to a search engine 216, and receiving one or more search results. Sadr: c.15 L.43-49. See 5_1 below).
While the combined teaching of Neilson, Paumard and Sadr does not explicitly teach, Gafni teaches: (5_1). obtaining within a graphical user interface a sketch of a portion of the scene (e.g., Methods that rely on text inputs only are more confined to generate within the training distribution, as demonstrated by [41]. Unusual objects and scenarios can be challenging to generate, as certain objects are strongly correlated with specific structures, such as cats with four legs, or cars with round wheels. The same is true for scenarios. “A mouse hunting a lion” is most likely not a scenario easily found within the dataset. By conditioning on scenes in the form of simple sketches, we are able to attend to these uncommon objects and scenarios, as demonstrated in Fig. 3, despite the fact that some objects do not exist as categories in our scene (mouse, lion). We solve the category gap by using categories that may be close in certain aspects (elephant instead of mouse, cat instead of lion). In practice, for non-existent categories, several categories could be used instead. Gafni: sec 4.7 para. 1 and Fig. 3);
(5_2). obtaining an image tile textual prompt associated with the sketch of the portion of the scene, wherein the image tile textual prompt specifies a desired characteristic of the respective image tile (e.g., We demonstrate the new capabilities this method provides in addition to controllability, such as (i) complex scene generation (Fig. 1), (ii) out-of-distribution generation (Fig. 3), (iii) scene editing (Fig. 4), and (iv) text editing with anchored scenes (Fig. 5). We additionally provide an example of harnessing controllability to assist with the creative process of storytelling in this video. Gafni: sec. 1 para. 9. Methods that rely on text inputs only are more confined to generate within the training distribution, as demonstrated by [41]. Unusual objects and scenarios can be challenging to generate, as certain objects are strongly correlated with specific structures, such as cats with four legs, or cars with round wheels. The same is true for scenarios. “A mouse hunting a lion” is most likely not a scenario easily found within the dataset. By conditioning on scenes in the form of simple sketches, we are able to attend to these uncommon objects and scenarios, as demonstrated in Fig. 3, despite the fact that some objects do not exist as categories in our scene (mouse, lion). We solve the category gap by using categories that may be close in certain aspects (elephant instead of mouse, cat instead of lion). In practice, for non-existent categories, several categories could be used instead. Gafni: sec 4.7 para. 1 and Fig. 3. It can be seen from Fig. 3 the non-existent scenarios: a mouse is sketched to hunt a lion, a car is sketched with triangular wheels and continuous tracks are sketched on a bicycle. It is obvious that combinations of effects of text editing and sketches can be conveniently be applied to generate non-existent scenarios);
Therefore it would have been obvious to a person of ordinary skill in the art before the effective filing date of the claimed invention to combine the teaching of Gafni into the combined teaching of Neilson, Paumard and Sadr so that uncommon objects and scenarios can be realized by sketching.
Regarding claims 14-15, the claims are non-transitory computer-readable media claims of method claims 4-5 respectively. The claims are similar in scope to claims 4-5 respectively and they are rejected under similar rationale as claims 4-5 respectively.
Allowable Subject Matter
Claims 6-10 and 16-19 are objected to being dependent upon rejected base claim. The claim would be allowable if rewritten to overcome the 35 U.S.C. 112(a) rejection of claims 1-20, and if rewritten in independent form including all the limitations of the base claim and any intervening claims.
The following is a statement of reasons for the indication of allowable subject matter in claim 6: The prior art of record, either individually or in combination, fails to teach the claimed limitation in the following:
obtaining within a graphical user interface a designation of a region of a corresponding image tile included in the plurality of image tiles;
obtaining a region textual prompt that corresponds to the region, wherein the region textual prompt specifies a desired characteristic of the region within the corresponding image tile; and
executing the machine learning model to generate the corresponding image tile based upon the region and the region textual prompt.
as recited in claim 6.
The following is a statement of reasons for the indication of allowable subject matter in claim 7: The prior art of record, either individually or in combination, fails to teach the claimed limitation in the following:
obtaining a selection of a previously generated image tile included in the plurality of image tiles;
generating a new image tile based upon a user input; obtaining a textual prompt specifying a desired characteristic for a combined image tile; and
executing the machine learning model to generate the combined image tile based upon the desired characteristic, the previously generated image tile, and the new image tile.
as recited in claim 7.
The following is a statement of reasons for the indication of allowable subject matter in claim 8: The prior art of record, either individually or in combination, fails to teach the claimed limitation in the following:
obtaining, within a graphical user interface, the spatial positioning of each image tile included in the plurality of image tiles relative to at least one other image tile included in the plurality of image tiles, wherein the spatial positioning comprises whitespace between at least a first image tile included in the plurality of image tiles and a second image tile included in the plurality of image tiles.
as recited in claim 8.
Regarding claims 9-10, the claims are dependent from claim 8 and they are objected under similar rationale as claim 8.
Claims 16-19 are similar in scope to claims 6-8 and 10 respectively and they are objected under similar rationale as claims 6-8 and 10 respectively.
Response to Arguments
Applicant's arguments filed 03/02/26 have been fully considered but they are not persuasive.
In regards to claim 1, the applicant argues that Neilson fails to teach scene composition. However, Neilson clearly provides composition of several scenes in Fig. 3a & Fig. 10 wherein the tiles are arranged to fit together to compose an accurate image. Therefore the applicant’s arguments in regards to claim 1 are unpersuasive in view of the teachings of Neilson.
Conclusion
Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a).
A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to Said Broome whose telephone number is (571)272-2931. The examiner can normally be reached Monday - Friday 8:30am-5pm.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/Said Broome/Supervisory Patent Examiner, Art Unit 2612