Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Claim Rejections - 35 USC § 102
The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action:
A person shall be entitled to a patent unless –
(a)(1) the claimed invention was patented, described in a printed publication, or in public use, on sale, or otherwise available to the public before the effective filing date of the claimed invention.
Claim(s) 21-23 and 26 is/are rejected under 35 U.S.C. 102(a)(1) as being anticipated by Shekhar et al. (U.S. Pub. 2017/0147906 A1).
With respect to claim 21, Sidhu discloses a method comprising:
obtaining a first record from a first source, the first record including a first text and a first image (i.e., “ From the plurality of unauthenticated pages, the online system filters out one or more unauthenticated pages that are associated with names of authenticated pages to obtain a group of candidate pages. Further, the online system pairs each candidate page up with an authenticated page. The candidate page has a name and/or image similar to the authenticated page. ”(abstract));
obtaining a second record from a second source, the second record including second text and a second image (i.e., “The online system maintains a plurality of authenticated pages and a plurality of unauthenticated pages, each of which is associated with a name and an image. From the plurality of unauthenticated pages, the online system filters out one or more unauthenticated pages that are associated with names of authenticated pages to obtain a group of candidate pages.”(abstract) and Examiner indicates second record is other authenticated pages and plurality of unauthenticated page and second source is other online system since the online system as define as reference is social networking system, have become increase prevalent in digital content distribution and consumption such as a celebrity or entity may create such a page to share information or some pape are created by a person or entity who pretends to be someone else (col. 1, lines 21-30));
determining that a textual similarity between the first text from the first record and the second text from the second record satisfies a first threshold (i.e., “ To determine a similarity score, in some embodiments, the pairing module 330 determines a name similarity score that indicates similarity between the name of the candidate page and the name of the authenticated page. ”(col. 10, lines 39-45));
comparing, in response to determining that the textual similarity satisfies the first threshold, the first image to the second image to determine an image similarity (i.e., “The module can also determine an image similarity score representing the similarity between images in the candidate and authentic pages. This scoring can be in response to the name similarity score being above a first similarity threshold, but can also be independent of that. In one embodiment, the pairing module 330 generates a hash value of the image of the candidate page and a hash value of the image of the authenticated page. ”(col. 10, lines 63-67) and step 450 in response step 440 at fig. 4);
determining, based at least in part on the textual similarity and the image similarity, that an overall similarity of the first record and the second record satisfies a second threshold (i.e., “the pairing module 330 can determine the overall similarity score based on the name similarity score, the image similarity score, or any combination of the two. There may also be more than one name or image similarity score for different key terms associated with the pages, for different images shown on the pages, etc. In addition, other features of the pages can be compared and scored, such as content in various posts displayed on the pages (e.g., text of the post, images of the post, likes or shares of the post by other users and which users, comments by other users on the post and which particular users), content in profiles of entities that operate the pages, etc… the overall similarity score equals the image similarity score or the name similarity score…the similarity score is an aggregation of the name similarity score and the image similarity score, or these scores plus any other scores calculated for the pages.”(col. 11, lines 15-25) and aggregation of name similarity score second threshold as claimed invention); and
generating, in response to determining that the overall similarity satisfies the second threshold, an association between the first record and the second record as corresponding to a same item (i.e., “The machine learning module 340 applies machine learning techniques to train the imposter detection model 350. When applied to a pair of candidate page and authenticated page, the imposter detection model 350 outputs an imposter score indicating a likelihood that the candidate page is an imposter page, i.e., whether the candidate page is intended for deceiving viewers that it is an authentic page of an individual or entity of the authenticated page. In one embodiment, the imposter score output from the imposter detection model 350 is a percentage from 0% to 100%. The output from the imposter detection model 350 can be in other forms.”(col. 11, lines 45-55)).
With respect to claim 22, Sidhu discloses wherein determining that the textual similarity satisfies the first threshold comprises: performing a field-by-field comparison of the first text and the second text such that first text from one field of the first record is compared to second text for a corresponding field of the second record (i.e., “ wherein determining a name similarity score comprises: comparing pronunciation of the name of the candidate page and the name of the authenticated page; and determining the name similarity score based on the comparison.”(claim 4)).
With respect to claim 23, Sidhu discloses wherein determining the image similarity comprises: determining image composition for the first image and the second image (i.e., “Taking the unauthenticated page 510 and authenticated pages 520 and 530 in FIG. 5 for example, the image filter 450 generates a semantic hashing code for each of the three pages based on visual features included in the images of the pages, including faces, clothes, background, hair, and hair accessory. ”(col. 14, lines 4-10) and image composition is visual feature included in the images of the pages such as faces, clothes and background, hair ); determining that a measure of similarity between the respective image compositions satisfies a third threshold (i.e., “ By comparing the semantic hashing code, the image filter 450 determines that compared with the image of the authenticated page 520, the image of the authenticated page 530 is more similar to the image of the unauthenticated page 510. Accordingly, the image similarity score of the authenticated page 530 is higher than that of the authenticated page 520. The image filter 450 pairs the unauthenticated page 510 with the authenticated page 530.”(col. 14, lines 4-10)); and comparing the similarly composed images to determine the image similarity threshold (i.e., “ By comparing the semantic hashing code, the image filter 450 determines that compared with the image of the authenticated page 520, the image of the authenticated page 530 is more similar to the image of the unauthenticated page 510. Accordingly, the image similarity score of the authenticated page 530 is higher than that of the authenticated page 520. The image filter 450 pairs the unauthenticated page 510 with the authenticated page 530.”(col. 14, lines 4-10));.
With respect to claim 26, Sidhu discloses wherein generating the association comprises generating one or more links between the first record and the second record, the generating one or more links comprising: generating a link between the first record and the second image of the second record (i.e., “The imposter detection module 230 may apply an initial filter to remove unauthenticated pages that are not intended for deceiving users. In some embodiments, the imposter detection module 230 filters out legitimate unauthenticated pages, the name of each of which is legitimately associated with the name of an authenticated page. For example, the imposter detection module 230 filters out an unauthenticated page whose name includes the name of an authenticated page and the word “fan” or “fans.” ”(col. 8, lines 24-32)); and generating a link between the second record and the first image of the first record (i.e., “installing an application associated with a content item, indicating a preference for a content item, sharing a content item with other users, interacting with an object associated with a content item, or performing any other suitable interaction. The online system 140 logs interactions between users presented with the content item or with objects associated with the content item. Additionally, the online system 140 receives compensation from a user associated with a content item as online system users perform interactions with a content item that satisfy the objective included in the content item.”(col. 5, lines 53-63)).
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claims 24-25 are rejected under 35 U.S.C 103 as being unpatentable over Sidhu (U.S. 11,032,316 B1) in view of Steruberg et al. (U.S. Pub. 2004/0133927 A1).
With respect to claim 24, Sidhu discloses all limitations recited in claim 23 except for determining the composition of the first image comprises: normalizing the first image; classifying the normalized first image; and segmenting the first image using one or more machine learning models to generate a segmentation mask indicating pixels of the first image corresponding to specific features.. However, Steruberg et al. discloses determining the composition of the first image comprises: normalizing the first image (i.e., “the value is selected to be the normalized intensity r of the red channel in the Digital Image (normalized to the interval 0 to 1), then the Warp Grid points will be seen to be attracted to the red areas of the Digital Image in the Warp process, the brightest red areas having the most attraction.”(0082) and “, meaning the transformed image pixel value is a function of the value of the corresponding pixel in the input image. Secondly, the transformed image may be a warp transform of the input image. The warp transform has been extensively discussed earlier in this application. Thirdly, the image transform may be a normalized two-dimensional DFT magnitude, directly analogous to the one-dimensional DFT discussed in the previous paragraph for audio input, and finally, the transformed image may be a histogram of the relative frequency of occurrence of identical m-by-n sub-images of the input image. ”(0621)); classifying the normalized first image (i.e., “ Thus the transformed data at 5001 would represent the normalized frequency of occurrence of each of the 4096 possible 3.times.4 binary sub-images. Each frequency of occurrence may be normalized by scaling by the inverse of the expected value of each sub-image frequency computed over all sub-images of all input images to the Digital Key database. Other methods of scaling are discussed later.”(0621)); and segmenting the first image using one or more machine learning models to generate a segmentation mask indicating pixels of the first image corresponding to specific features. (i.e.,” The process does not provide any obvious means for subdividing the database into smaller segments, one of which can be known a priori to contain the unknown image. Thus, the computer performing the comparisons must do what a human would have to do: compare each database image to the unknown image one at a time on a pixel-by-pixel basis. Even for a high-speed computer, this is a very time consuming process.” (0042) and “The term pattern recognition has come to represent all such methods. Examples of such feature sets, which can be extracted and used, might be line segments, defined, perhaps, by the locations of the endpoints, by their orientation, by their curvature, etc. The reduction of images to feature sets is always an attempt to translate image composition, for which, there is no language, into a restrictive dictionary of image features”(0044)). It would have been obvious for a person of ordinary skill in the art, before the effective filing date of the claimed invention, to include Sternberg et al.’s feature in order to have accurate in recognize the image for the stated purpose has been well known in the art as evidenced by teaching of Sternberg et al. (0002)
With respect to claim 24, Sternberg discloses the method of claim 24, wherein comparing the similarly composed images to determine the image similarity comprises: using the segmentation mask to determine a first portion of the first image corresponding to a representation of an item (i.e., “ Here, the image referred to is the transformed image of the auxiliary construct previously discussed. The reticle mask 6002 is a complex spatial filter whose individual elements weight the transmitted light rays 6003 from the display by +1 or -1, the -1 weighting being accomplished by 180 degree phase shifting of the ray 6003’(0633)); and comparing the first portion of the first image with a corresponding second portion of the second image (i.e., “If we are required to find the best matching Visual Key Vector from a database, each Visual Key Vector in the database will need to be compared to the corresponding vector of the Query Visual Key Vector. For the example in the preceding paragraph, that would represent 16*16*2*1 million of 8-byte comparisons. If each 8-byte comparison took ten nanoseconds (10.sup.-8 seconds) then a best match search of the database of 1 million would take 5.12 seconds, disregarding any other necessary computations required for determining the match distance”(0420) and “From the preceding discussions it can be concluded that a best match to a given Visual Key Vector may be obtained by pairwise comparing the given Visual Key Vector to all the Visual Key Vectors in the database and noting which one yields the closest match. The question of whether a given database contains a match to a given Visual Key Vector is equivalent to the question of whether the best match in a database is sufficiently close to be considered to have arisen from the same Picture. Thus the matching distance of the best match must be compared to a specified maximum allowable matching distance to be considered to have arisen from the comparing of Visual Key Vectors derived from the same Picture”(0415)).
Claim 27 is rejected under 35 U.S.C 103 as being unpatentable over Sidhu (U.S. 11,032,316 B1) in view of Shekhar et al. (U.S. Pub. 2017/0147906 A1).
With respect to claim 27, Sidhu discloses all limitations recited in claim 23 except for normalizing text from one or more of the first record and the second record prior to determining the textual similarity between the first text and the second text. However, Shekhar et al. discloses normalizing text from one or more of the first record and the second record prior to determining the textual similarity between the first text and the second text (i.e., “The process for determining 212 a normalized text feature score 256 begins with the extraction of at least one text feature from the text of the video”(0025) and “Once the recursive autoencoder has analyzed 244 the text and extracted a semantic vector from a text feature, and once the deep neural network learning analysis 232 has identified semantic descriptions of objects in a video feature, these two semantic meaning are compared to determine text/image meaning
similarity 248. ”(0026) and fig. 2 shows step 212 prior step 232). )). It would have been obvious for a person of ordinary skill in the art, before the effective filing date of the claimed invention, to include Shekhar et al.’s feature in order to improve for editor and easy to edit video for the stated purpose has been well known in the art as evidenced by teaching of Sternberg et al. (0002)
With respect to claims 28-40, the claims 28-40 are rejected as set of claims 21-27 since the claims 28-40 are similar set of claims 21-27 but different form.
Double Patenting
The nonstatutory double patenting rejection is based on a judicially created doctrine grounded in public policy (a policy reflected in the statute) so as to prevent the unjustified or improper timewise extension of the "right to exclude" granted by a patent and to prevent possible harassment by multiple assignees. A nonstatutory obviousness-type double patenting rejection is appropriate where the conflicting claims are not identical, but at least one examined application claim is not patentably distinct from the reference claim(s) because the examined application claim is either anticipated by, or would have been obvious over, the reference claim(s). See, e.g., In re Berg, 140 F.3d 1428, 46 USPQ2d 1226 (Fed. Cir. 1998); In re Goodman, 11 F.3d 1046, 29 USPQ2d 2010 (Fed. Cir. 1993); In re Longi, 759 F.2d 887, 225 USPQ 645 (Fed. Cir. 1985); In re Van Omum, 686 F.2d 937, 214 USPQ 761 (CCPA 1982); In re Vogel, 422 F.2d 438, 164 USPQ 619 (CCPA 1970); and In re Thorington, 418 F.2d 528, 163 USPQ 644 (CCPA 1969).
A timely filed terminal disclaimer in compliance with 37 CFR 1.321(c) or 1.321(d) may be used to overcome an actual or provisional rejection based on a nonstatutory double patenting ground provided the conflicting application or patent either is shown to be commonly owned with this application, or claims an invention made as a result of activities undertaken within the scope of a joint research agreement. Effective January 1, 1994, a registered attorney or agent of record may sign a terminal disclaimer. A terminal disclaimer signed by the assignee must fully comply with 37 CFR 3.73(b).
Claims 21-40 are rejected on the ground of nonstatutory obviousness-type double patenting as being unpatentable over claims 1-20 of Patent No. 12,423,361. Although the conflicting are not patentably distinct from each other because since the claims of the Patent No. 12,423,361 contains every element of the claims of the instant application, and as such, anticipate the claims of the instant application. (see table below).
Instant Application claim 1
Patent No. . 12,423,361 claim 1
A method comprising:
obtaining a first record from a first source, the first record including a first text and a first image;
obtaining a second record from a second source, the second record including second text and a second image;
determining that a textual similarity between the first text from the first record and the second text from the second record satisfies a first threshold;
comparing, in response to determining that the textual similarity satisfies the first threshold, the first image to the second image to determine an image similarity;
determining, based at least in part on the textual similarity and the image similarity, that an overall similarity of the first record and the second record satisfies a second threshold; and generating, in response to determining that the overall similarity satisfies the second threshold, an association between the first record and the second record as corresponding to a same item.
A method for web crawling comprising: retrieving, by a computer system, a webpage including instructions for rendering a plurality of elements; ingesting, by a crawling engine executing on the computer system, a set of extraction instructions;
traversing, by the crawling engine, the plurality of elements of the webpage to identify selected elements from the plurality of elements based at least in part on the set of extraction instructions;
applying, by the crawling engine, the set of extraction instructions to the selected elements of the webpage to extract selected data of the webpage according to the set of extraction instructions, wherein the selected data includes at least a first product image;
extracting, by the crawling engine, second extracted data associated with a second webpage having a different domain than the webpage and not referenced in the webpage, wherein the second selected data includes at least a second product image;
determining that a textual similarity between the selected data from the webpage and the second extracted data associated with the second webpage satisfies a first threshold;
comparing, in response to determining that the textual similarity satisfies the first threshold and by the computer system, the first product image to the second product image to determine an image similarity;
determining, by the computer system and based at least in part on the textual similarity and the image similarity, that the similarity of the webpage and the second webpage satisfies a second threshold;
generating, in response to determining that the similarities satisfy the second threshold and by the computer system, an association between the webpage and the second webpage as corresponding to the same product;
determining textual information included in the second extracted data associated with the second webpage and not included in the selected data of the webpage; and adding at least a portion of the textual information to content from the webpage.
Claim Rejections - 35 USC § 101
35 U.S.C. 101 reads as follows:
Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title.
Claims 1-21 are rejected under 35 U.S.C. 101 because:
At step 1:
Claims 21-40 are directed to a “Data extraction approach for retail crawling engine” and thus directed to a statutory category.
At step 2A, Prong One:
The claim recites the following limitation directed to an abstract ideas:
“obtaining a first record from a first source, the first record including a first text and a first image” recites a mental process as obtaining a first record from a first source, the first record including a first text and a first image.
. “obtaining a second record from a second source, the second record including second text and a second image” recites a mental process as obtaining a second record from a second source, the second record including second text and a second image;
“determining that a textual similarity between the first text from the first record and the second text from the second record satisfies a first threshold” recites the mental process determining that a textual similarity between the first text from the first record and the second text from the second record satisfies a first threshold.
Father the claim 21-27 are directed to abstract ideas
Claim 22 recites wherein determining that the textual similarity satisfies the first threshold comprises: performing a field-by-field comparison of the first text and the second text such that first text from one field of the first record is compared to second text for a corresponding field of the second record.. It is abstract idea, and it is mental process.
Claim 23 recites determining the image similarity comprises: determining image composition for the first image and the second image; determining that a measure of similarity between the respective image compositions. It is abstract idea, and it is mental process.
Claims 24 recites wherein determining the composition of the first image comprises: normalizing the first image; classifying the normalized first image; and segmenting the first image using one or more machine learning models to generate a segmentation mask indicating pixels of the first image corresponding to specific features. It is abstract idea, and it is mental process.
Claim 25 recites comparing the similarly composed images to determine the image similarity comprises: using the segmentation mask to determine a first portion of the first image corresponding to a representation of an item; and comparing the first portion of the first image with a corresponding second portion of the second image. It is abstract idea, and it is mental process.
Claim 26 recites generating the association comprises generating one or more links between the first record and the second record, the generating one or more links comprising: generating a link between the first record and the second image of the second record; and generating a link between the second record and the first image of the first record. It is abstract idea, and it is mental process.
Claim 27 recites normalizing text from one or more of the first record and the second record prior to determining the textual similarity between the first text and the second text. It is abstract idea, and it is mental process.
At step 2A, Prong Two:
The claims recite the following additional elements:
That the content management system includes “server” “resource”, which are high level recitation of generic computer component s and functions and represent mere instruction to apply to a computer as in MPEP 2106.05 (f) which does not provide integration into a practical application.
At step 2B
The conclusions for the mere implementation using a generic computer and mere field of use are carried over and to not provide significantly more.
With respect to claims 28-40, the claims 28-40 are rejected as set of claims 21-27 since the claims 28-40 are similar set of claims 21-27 but different form.
Conclusion
Any inquiry concerning this communication or earlier communications from the examiner should be directed to HUNG T VY whose telephone number is (571)272-1954. The examiner can normally be reached M-F 8-5.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Tony Mahmoudi can be reached at (571)272-4078. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/HUNG T VY/Primary Examiner, Art Unit 2163 August 12, 2026