DETAILED ACTION
1. Applicant’s election without traverse of Group III (claims 17-20) in the reply filed on 06/18/2026 is acknowledged. Claim 17 is independent.
Notice of Pre-AIA or AIA Status
2. The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Priority
3. This application, filed on 10/08/2024 doesn’t claim priority.
Information Disclosure Statement
4. The information disclosure statements (IDS), filed on 11/29/2024; 03/29/2025; 02/26/2026 and 06/30/2026 have been considered. The submission is in compliance with the provisions of 37 CFR 1.97. Form PTO-1449 is signed and attached hereto.
Drawings
5. The drawings filed on October 8, 2024 are accepted.
Specification
6. The specification filed on October 8, 2024 is also accepted.
Internet Communications
Applicant is encouraged to submit a written authorization for Internet communications (PTO/SB/439, http:/www.uspto.gov/sites/default/files/documents/sb0439.pdf) in the instant patent application to authorize the examiner to communicate with the applicant via email. The authorization will allow the examiner to better practice compact prosecution. The written authorization can be submitted via one of the following methods only: (1) Central Fax, which can be found in the Conclusion section of this Office action; (2) regular postal mail; (3) EFS WEB; or (4) the service window on the Alexandria campus. EFS web is the recommended way to submit the form since this allows the form to be entered into the file wrapper within the same day (system dependent). Written authorization submitted via other methods, such as direct fax to the examiner or email, will not be accepted. See MPEP § 502.03.
Claim Rejections - 35 USC § 101
8. 35 U.S.C. 101 reads as follows:
Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title.
Claims 17-20 are rejected under 35 U.S.C. 101 because the claimed invention is directed to non-statutory subject matter.
Regarding independent Claim 17 and dependent claims 18-20, the preamble recites a “one or more computer-readable storage media”. Pending claims are interpreted as broadly as their terms reasonably allow. The broadest reasonable interpretation of a claim drawn to a computer readable storage medium (also called machine readable medium and other such variations) typically covers forms of non-transitory tangible media and transitory propagating signals per se in view of the ordinary and customary meaning of computer readable media (See MPEP 2111.01). When the broadest reasonable interpretation of a claim covers a signal per se, the claim must be rejected under 35 U.S.C. §101 as covering non-statutory subject matter. Therefore, the instant Claims 17-20 do not fall within at least one of the four categories of patent eligible subject matter because under the broadest reasonable interpretation a “computer-readable storage medium” can encompass non-statutory transitory forms of signal transmission such as a propagating electrical or electromagnetic signal per se, and therefore is directed to non-statutory subject matter.
Examiner suggests that a claim drawn to such a computer readable medium that covers both transitory and non-transitory embodiments may be amended to narrow the claim, cover only statutory embodiments, and thus avoid rejection under 35 U.S.C. §101, by adding the limitation "non-transitory" to the claim [or any similar limitations such as "computer usable memory", or "computer usable storage memory", or "computer readable memory", or "computer readable device", (i.e. any variations thereof, where "media" or "medium" is replaced by "device" or "memory") ]. Such an amendment would typically not raise the issue of new matter, even when the specification is silent because the broadest reasonable interpretation relies on the ordinary and customary meaning that includes signals per se. The limited situations in which such an amendment could raise issues of new matter occur, for example, when the specification does not support a non-transitory embodiment because a signal per se is the only viable embodiment such that the amended claim is impermissibly broadened beyond the supporting disclosure. See, e.g., Gentqv Galleiy, Inc. v. Berkline Corp., 134 F.3d 1473 (Fed. Cir. 1998).
Claims 17-20 are directed to a judicial exception, namely abstract idea, without significantly more. The claim(s) recite(s) receiving information, evaluating correspondence between membership identifiers, grouping dataset records, generating probabilistic summaries from identity keys and attributes and storing the summaries to provide a probabilistic result of a query. Collectively, these limitations recite organizing and evaluating information which falls within the mental-process grouping of abstract ideas and generating a probabilistic representation or result, which falls within the mathematical-concept grouping of abstract ideas. These judicial exceptions are not integrated into a practical application and the claim(s) does/do not include additional elements that are sufficient to amount to significantly more than the judicial exception because the addition elements such as one or more computer readable storage media, a processing derive and a database-do not integrate the abstract idea into a practical application. These elements merely provide a computer environment in which the information is received, grouped, summarized, and stored. The claims do not recite a particular sketch-generation algorithm, hashing arrangement, database architecture, memory configuration or other specific mechanism that changes the operations of the processing device or the database. Instead, the claim recites the desired results of producing probabilistic sketches and using the sketches to answer a query.
Claim Rejections - 35 USC § 103
10. The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
(a) A patent may not be obtained though the invention is not identically disclosed or described as set forth in section 102 of this title, if the differences between the subject matter sought to be patented and the prior art are such that the subject matter as a whole would have been obvious at the time the invention was made to a person having ordinary skill in the art to which said subject matter pertains. Patentability shall not be negatived by the manner in which the invention was made.
11. Claims 17-19 are rejected under 35 U.S.C. 103 as being unpatentable over Ronen Cohen (herein after referred as Cohen) (NPL document titled, “Applied Probability - Counting Large Set of Unstructured Events with Theta Sketches”) (June 29, 2020) in view of Craig Wright et al (Wright) (US Publication No. 20220376887 A1 (Nov 24, 2022)
12. The factual inquiries set forth in Graham v. John Deere Co., 383 U.S. 1, 148 USPQ 459 (1966), that are applied for establishing a background for determining obviousness under 35 U.S.C. 103(a) are summarized as follows:
1. Determining the scope and contents of the prior art.
2. Ascertaining the differences between the prior art and the claims at issue.
3. Resolving the level of ordinary skill in the pertinent art.
4. Considering objective evidence present in the application indicating obviousness or non-obviousness.
The following is referring to independent claim 17,
As per independent claim 17, Cohen
receiving a structured dataset including a plurality of dataset records that include a respective identity key of a plurality of identity keys, a plurality of attributes associated, respectively, with the plurality of identity keys, and a plurality of membership identifiers associated with respective said attributes [See “The data model, para. 2”, “Events sent to AppsFlyer carry common information such as the event’s timestamp, the internal app-id, the non-PII user ID” and “event_attributes - a dictionary of key-value pairs” and see The data model, para. 9, Unified model includes , “event-type”; “event_attributes” and “Original fields including app_id and timestamp” event records corresponds to the limitation “dataset records”, “the user_id or non-PII user ID” corresponds to the limitation, “identity key”; the event_attributes” corresponds to the limitation, “attributes” The app identifier, event type, event date and event attribute values are membership-identifying fields that determine which records belong to later -formed groups]
forming a plurality of dataset groups by grouping the dataset records based on correspondence with membership identifiers of the plurality of membership identifiers [The persistence System, para. 1, “all events belonging to the same combination of app_id, event type and event attributes are grouped together”, The expressly groups records by corresponding values of app identifier, event type and event attributes, These are the membership identifiers because they identify the records’ membership in a particular dataset group];
generating a plurality of sketches, respectively, based on the plurality of dataset groups [The persistence System, para. 1, “their user IDs are aggregated into a single Theta Sketch” Cohen generates a sketch for each grouped combination of event criteria. Because there are many combinations of app identifier, event type, and event attributes, Cohen teaches a plurality of sketches respectively based on the plurality of dataset groups]
each said sketch configured as a probabilistic data structure based on one or more said identity keys and one or more said attributes of the plurality of dataset records associated with a respective said group [Why Theta Sketches, para. 1, “Theta Sketches were selected, a probabilistic data structure” and Why Theta Sketches, para. 1, sketch supports “getting back an approximation”, The Persistence System, grouped events have “their user IDs …aggregated into a single Theta Sketch” Cohen’s Theta Sketch is expressly a probabilistic data structure. It is based on user indenters from records that are grouped according to app identifier, event type, and event attributes, so the sketch is based on identity keys and attributes associated with the respective group.];and
storing the plurality of sketches in a database that supports a probabilistic result to a query [The HBase model, 1st para. “Each row in this table corresponds to some combination of app_id, event date, event name and event attributes, and contains the serialized pre-aggregated”; a query finds “all precomputed Theta Sketches in Hbase” and the scan “will return multiple rows“ whose sketches are used for the result, The qurery model, para. 5, The system calculates an “estimated” audience size rather than an exact one. Thus, Cohen stores sterilized pre-aggregated sketches in an HBase database. Queries retrieve relevant stored sketches and combine them to produce estimated audience-size results, which are probabilistic results rather than exact row-level results].
Cohen doesn’t explicitly disclose the following underlined claim limitation “one or more computer-readable storage media storing instructions that, responsive to execution by a processing device, causes the
However, Wright discloses the above underlined limitation: “one or more computer-readable storage media storing instructions that, responsive to execution by a processing device, causes the[Para. 0179 ” The memory stores machine instructions that, when executed by processor, cause processor to perform one or more of the operations described herein“]
Cohen and Wright are in the same field of endeavor and are directed to maintaining and accessing a set of data records including identifiers and attributes.
It would have been obvious to one having ordinary skill in the art, before the effective filing of the claimed invention, to modify the system of Cohen, using a mechanism such as “one or more computer-readable storage media storing instructions that, responsive to execution by a processing device, causes the” as taught by Wright because this would enhance the security and privacy of data and data management systems through storage and the use of secured probabilistic data structure [See Wright, para. 0007, “the systems and methods discussed herein can provide increased security and privacy of data and data counting systems through the use of encrypted probabilistic data structures and a homomorphic encryption scheme”]
The following is referring to dependent claims 18-19
As per dependent claim 18, the combination of Cohen and Wright discloses the method or the computer readable medium as applied to claim 17 above. Furthermore, Cohen discloses the method, wherein the plurality of sketches is stored independent of row-level data of the plurality of membership identifiers and do not support direct identification of the plurality of membership identifiers via the database [The HBase model, 1st para. “Each row in this table corresponds to some combination of app_id, event date, event name and event attributes, and contains the serialized pre-aggregated”; a query finds “all precomputed Theta Sketches in Hbase” and the scan “will return multiple rows“ whose sketches are used for the result, The qurery model, para. 5, The system calculates an “estimated” audience size rather than an exact one. Thus, Cohen teaches storing “serialized pre-aggregated Theta Sketch” rows rather than raw event rows, and Cohen explains that Theta Sketches retain only a sample and “cannot be used to test for the existence of an element” That supports storing sketches independent of row-level membership data and not directly identifying membership identifiers through the database].
As per dependent claim 19, the combination of Cohen and Wright discloses the method or the computer readable medium as applied to claim 17 above. Furthermore, Wright discloses the method/system/computer readable media, wherein, wherein the structured dataset is a structured customer dataset having the plurality of membership identifiers as confidential information and wherein the plurality of sketches do not include the confidential information as stored in the database [Cohen’s system uses non-personally identifying user identifier and stores sketches and Wright furthermore on Para. 0006 discloses that, information that can include private or protected information and on at least para. 0008, describes producing estimations, “without transmitting protected or private information”]
13. Claim 20 is rejected under 35 U.S.C. 103 as being unpatentable over Ronen Cohen (herein after referred as Cohen) (NPL document titled, “Applied Probability - Counting Large Set of Unstructured Events with Theta Sketches”) (June 29, 2020) in view of Craig Wright et al (Wright) (US Publication No. 20220376887 A1 (Nov 24, 2022) and further in view of Alberton et al (Alberton)(US Publication No. 20180268166 A1, Pub. Date: Sept. 20, 2018)
As per dependent claim 20, the combination of Cohen and Wright discloses the method or the computer readable medium as applied to claim 17 above. Furthermore, Cohen discloses wherein the generating is based on the filtering [See figure 7, Query on Event table, “filters”]
The combination of Cohen and Wright doesn’t disclose the following underlined claim limitation:
the operations further comprise filtering confidential information from the structured dataset by removing reference to the plurality of membership identifiers from the structured dataset.
However, Alberton discloses: “filtering confidential information from the structured dataset by removing reference to the plurality of membership identifiers from the structured dataset [Para. 0102, “removing all user-identity from the events and user attributes before they are passed to the content processing component” and para.0104 explains that each event contains a user identifier and that the original identifier are replaced with anonymized tokens and para. 0105 further explain that the original user identifier are not accessible to the downstream content processing stage, The original user identifier corresponds to membership identifier because it identifies the particular platform member associated with the event and its attributes. Removing the link between the event or attribute and that underlying member therefore filters confidential identity information from the dataset]
It would have been obvious to one having ordinary skill in the art, before the effective filing of the claimed invention, to modify the system of Cohen and Wright, using a mechanism such as “filtering confidential information from the structured dataset by removing reference to the plurality of membership identifiers from the structured dataset” as taught by Alberton because this would enhance the privacy and security of the system by preserving user privacy and ensure that individual users can never be identified from the aggregate information released by the system and prevent the release of any information that could be attributed to individual users.[ See Alberton para. 0007, ” In order to preserve user privacy,…The goal is to ensure that individual users can never be identified from the aggregate information released by the system and prevent the release of any information that could be attributed to individual users].
Conclusion
14. The prior art made of record and not relied upon is considered pertinent to applicant's disclosure.
US Pub. No. 20220292101A1 Mai [Same Assignee] discloses the method for responding to a query requesting an intersection being performed. The method includes receiving a query referencing a first set, a second set, and a desired quantile related to the first set from among a plurality of quantiles; generating a data structure including a bottom-k sketch of user identifiers (ids) of the first set and corresponding numerical values of the first data; partitioning the data structure into a plurality of sketches to correspond to the quantiles, respectively; determining an intersection of one of the sketches associated with the desired quantile and a sketch of the second set; and responding to the query based on the intersection.
US Pub. No. 20210357403A1 Dash discloses methods for distributed histogram computation in a framework utilizing data stream sketches and samples are performed by systems and devices. Distributions of large data sets are scanned once and processed by a computing pool, without sorting, to generate local sketches and value samples of each distribution. The local sketches and samples are utilized to construct local histograms on which cardinality estimates are obtained for query plan generation of distributed queries against distributions. Local statistics of distributions are also merged and consolidated to construct a global histogram representative of the entire data set. The global histogram is utilized to determine a cardinality estimation for query plan generation of incoming queries against the entire data set. The addition of new data to a data set or distribution involves a scan of the new data from which new statistics are generated and then merged with existing statistics for a new global histogram.
NPL Document, Titled “Sketch Techniques for Approximate Query Processing” by Graham Cormode discloses Sketches, are designed so that the update caused by each new piece of data is largely independent of the current state of the summary. This design choice makes them faster to process, and also easy to parallelize. “Frequency based sketches” are concerned with summarizing the observed frequency distribution of a dataset. From these sketches, accurate estimations of individual frequencies can be extracted. This leads to algorithms to find the approximate heavy hitters (items which account for a large fraction of the frequency mass) and quantiles (the median and its generalizations).
US Pub. No. 20220036390A1 Sheppard discloses methods, apparatus, systems, and articles of manufacture to estimate audience measurement metrics based on users represented in Bloom filter arrays are disclosed. An apparatus includes a communications interface to receive a first Bloom filter array from a first computer of a first database proprietor. The first Bloom filter array is representative of first users who accessed media. The first users are registered with the first database proprietor. The first Bloom filter array includes a first array of first elements. Values of respective ones of the first elements are either a 0 or a 1 based on whether quantities of the first users allocated to the respective ones of the first elements are even or odd. The apparatus further includes a Bloom filter array analyzer to estimate a first cardinality for the first Bloom filter array. The first cardinality is indicative of a total number of the first users who accessed the media.
US Pub. No. 20140280153A1 Cronin discloses methods for implementing a GROUP command with a predictive query interface including means for generating indices from a dataset of columns and rows, the indices representing probabilistic relationships between the rows and the columns of the dataset; storing the indices within a database of a host organization; exposing the database of the host organization via a request interface; receiving, at the request interface, a query for the database specifying a GROUP command term and a specified column as a parameter for the GROUP command term; querying the database using the GROUP command term and passing the specified column to generate a predictive record set; and returning the predictive record set responsive to the query, the predictive record set having a plurality of groups specified therein, each of the returned groups of the predictive record set including a group of one or more rows of the dataset. Other related embodiments are further disclosed.
US Pub. No. 20240045866A1 Bordawekar discloses methods or computer program products to facilitate receiving results of a semantic structured query language (SQL) query and employing sparse hash-table based sketches to interpret a semantic structured query language (SQL) query result. A computing component stores a first space-efficient structure sketch in a compressed serialize form. The computing component can load a second space-efficient data structure sketch along with the first space-efficient data structure sketch and can compute one or more interpretability scores by extracting co-occurrence information from the first space-efficient data structure sketch. The second space-efficient data structure sketch can include a sketch for containment check.
See the other cited prior art/s
Any inquiry concerning this communication or earlier communications from the examiner should be directed to SAMSON B LEMMA whose telephone number is 571-272-3806. The examiner can normally be reached on M-F 8am-10pm.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Shaw Yin Chen can be reached on to 571-272-8878. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of an application may be obtained from the Patent Application Information Retrieval (PAIR) system. Status information for published applications may be obtained from either Private PAIR or Public PAIR. Status information for unpublished applications is available through Private PAIR only. For more information about the PAIR system, see http://pair-direct.uspto.gov. Should you have questions on access to the Private PAIR system, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative or access to the automated information system, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/SAMSON B LEMMA/Primary Examiner, Art Unit 2498