DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Applicant’s arguments and amendments filed 7/15/2026 have been entered and considered but are not completely persuasive.
The Terminal Disclaimers filed 7/15/2026 have been entered, obviating the rejections under non-statutory double patenting.
The IDS filed 7/15/2026 has been entered and considered.
Claims 1-13 are under examination.
Claim Interpretation
The claims in this application are given their broadest reasonable interpretation (BRI) using the plain meaning of the claim language in light of the specification as it would be understood by one of ordinary skill in the art.
Claim Rejections - 35 USC § 101
35 U.S.C. 101 reads as follows:
Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title.
Claims 1-13 remain rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea of mental steps, mathematic concepts, organizing human activity, or a natural law without significantly more.
Applicant is directed to MPEP 2106 for the most current and complete guidelines in the analysis of patent- eligible subject matter. The current MPEP is the primary source for the USPTO’s patent eligibility guidance.
With respect to step (1): YES, the claims are drawn to statutory categories: Claims 1-11 are directed to computer-implemented methods. Claim 12 is drawn to a system for carrying out the method. Claim 13 is drawn to a non-transitory computer readable media comprising instructions for carrying out the method.
With respect to step (2A) (1): YES, the claims recite an abstract idea, law of nature and/or natural phenomenon (MPEP 2106.04(a)(2)). The claims explicitly recite elements that, individually and in combination, constitute one or more judicial exceptions (JE).
Mathematic concepts, Mental Processes or Elements in Addition (EIA) in the claim(s) include:
Claim 1. A method for identifying a conserved target region in a subset of a plurality of nucleic acid sequences, comprising:
at a system comprising one or more processors and memory storing instructions executable by the processor:
(Preamble, reciting a method, and EIA the use of a general-purpose computer system. MPEP 2106.05(h).)
receiving genomic data representing a plurality of nucleic acid sequences;
(EIA- data gathering step of receipt of nucleic acid sequence data. MPEP 2106.05(d, g))
selecting a subset of the plurality of nucleic acid sequences, wherein the nucleic acid sequences in the subset share a characteristic;
(Mental Process of observing the plurality, observing a shared unidentified characteristic of a subset, and making a judgement as to what subset to select. MPEP 2106.04(a)(2).)
creating and storing data in a first index representing a first set of the plurality of nucleic acid sequences corresponding to the subset, wherein the first index comprises at least 412 elements representing each respective permutation of nucleic acid sequences, and
wherein the data created and stored in the first index comprises a first plurality of data structures each associated with a respective nucleic acid sequence of the first set;
creating and storing data in a second index representing a second set of the plurality of nucleic acid sequences distinct from the subset, wherein the second index comprises at least 412 elements representing each respective permutation of nucleic acid sequences, and
wherein the data created and stored in the second index comprises a second plurality of data structures each associated with a respective nucleic acid sequence of the second set; and
(Mathematic Concept: Creating permutations of a nucleic acid sequence requires combinatorial mathematics that calculate and assign integer values to a specific sequence of nucleotides of length k. The permutation index converts the text string of the k-mer into a distinct numerical rank, uniquely mapping it to a single value within a space of 4^{k} total possible permutations. MPEP 2106.04(a)(2). The “index” merely stores the data and has no additional structural requirements or functions. The “associated data structures” have no particular structure or functional aspects. MPEP 2106.04(a)(2))
identifying a target region that is conserved in the subset of nucleic acids, wherein the identifying comprises:
identifying, by the first index, a region of the target region as a conserved region appearing in every nucleic acid sequence in the first set;
(Mental step of observing a region present in every sequence of the first subset, which requires observation, matching, and making a judgement as to whether any conserved region exists. [0023] “conserved (e.g. matching, identical)” MPEP 2106.04(a)(2))
determining, by the second index, that the conserved region appears in none of the nucleic acid sequences in the second set; and
(Mental step of observing the conserved region is not present in the second subset, which requires observation, matching, and making a judgement as to whether the conserved region exists in the second subset. [0023] “conserved (e.g. matching, identical)” MPEP 2106.04(a)(2))
generating and outputting data representing the identified conserved target region.
(EIA- routine data output MPEP 2106.05(b))
2. (Previously Presented) The method of claim 1, wherein creating and storing data in the first index comprises: for each of the nucleic acid sequences in the first set, dividing the nucleic acid sequence into a plurality of sub-strings; for each of the plurality of sub-strings, storing a respective one of the first plurality of data structures in the first index, wherein the respective one of the first plurality of data structures indicates an identity of the nucleic acid sequence, a permutation of bases forming the sub-string, and a position of the sub-string in the nucleic acid sequence.
(Mathematic Concept: Creating permutations of a nucleic acid sequence requires combinatorial mathematics that calculate and assign integer values to a specific sequence of nucleotides of length k. The permutation index converts the text string of the k-mer into a distinct numerical rank, uniquely mapping it to a single value within a space of 4^{k} total possible permutations. MPEP 2106.04(a)(2). The “index” merely stores the data and has no additional structural requirements or functions. The “associated data structures” have no particular structure or functional aspects. MPEP 2106.04(a)(2))
3. (Currently Amended) The method of claim 2, wherein identifying the conserved region appearing in every nucleic acid sequence in the first set comprises determining, for a given sub- string of a first nucleic acid sequence of the first set, that a corresponding first data structure stored in the first index indicates a common permutation of bases as a second data structure stored in the first index for a second nucleic acid sequence in the first set.
(Mental process of observing information about “a given sub-string”)
4. (Currently Amended) The method of claim 3, wherein identifying the conserved region appearing in every nucleic acid sequence in the first set comprises determining that the second data structure indicates: an identity for the second nucleic acid sequence that matches an identity of a nucleic acid sequence that has been determined to include a previously-matched sub-string, wherein the previously-matched sub-string matches the first nucleic acid sequence at a span occurring immediately before the given sub-string in the first nucleic acid sequence; and a position in the second nucleic acid sequence corresponding to a span occurring immediately after the previously-matched sub-string.
(Mental processes of observation of a match, the location of the match, and the position of the match.)
5. (Previously Presented) The method of claim 3, wherein the determination is performed iteratively with respect to different sub-strings of the first nucleic acid sequence and different data structures in the first index, until a plurality of adjacent sub-strings of the first nucleic acid sequence are determined to occur in a same order in each of the other nucleic acid sequences in the first set,wherein the plurality of adjacent sub-strings of the first nucleic acid sequence together are at least a predefined minimum number of bases in length.
(Mental process of repeating the determination of the previous step, for different substrings, and mathematic concept of counting the adjacent substring nucleic acids to meet a minimum length requirement.)
6. (Currently Amended) The method of claim 2, wherein confirming that the region determining that the conserved region appears in none of the nucleic acid sequences in the second set comprises: determining, for at least one given sub-string of a nucleic acid sequence of the first set, whether a data structure stored in the second index for a nucleic acid sequence in the second set indicates all three of: a common permutation of bases as indicated by a data structure stored in the first index for the nucleic acid sequence of the first set; an identity for the nucleic acid sequence of the second set that matches an identity of a nucleic acid sequence that has been determined to include a previously-matched sub-string, wherein the previously-matched sub-string matches the nucleic acid sequence of the first set at a span occurring immediately before the given sub-string in the nucleic acid sequence of the first set; and a position in the nucleic acid sequence of the second set corresponding to a span occurring immediately after the previously-matched sub-string.
(Mental processes of comparing one substring of the first set, to data present in the second index, observing the results, and making a judgement as to whether the three conditions are met.)
7. (Previously Presented) The method of claim 6, wherein the determination is performed iteratively with respect to different sub-strings of the nucleic acid sequence of the first set in order to determine that, for every nucleic acid sequence in the second index, at least one data structure fails at least one of the three conditions for at least one sub-string in the region of the first nucleic acid sequence.
(Mental process of performing the steps of claim 6 iteratively, with different substrings, comparing one substring of the first set, to data present in the second index, observing the results, and making a judgement as to whether one of the three conditions are met.)
8. (Previously Presented) The method of claim 1, wherein the plurality of nucleic acid sequences comprises one of DNA, cDNA, RNA, mRNA, PNA.
(EIA- data gathering limitation, describing an aspect of the data gathered. MPEP 2106.05(g))
9. (Previously Presented) The method of claim 1, wherein creating and storing data in the second index comprises: for each of the nucleic acid sequences in the second set, dividing the nucleic acid sequence into a plurality of sub-strings; for each of the plurality of sub-strings, storing a respective one of the second plurality of data structures in the second index, wherein the respective one of the second plurality of data structures indicates an identity of the nucleic acid sequence, a permutation of bases forming the sub-string, and a position of the sub-string in the nucleic acid sequence.
(EIA and Mental Process and Mathematic Concept: The creation of the index requires “dividing the nucleic acid sequence into substrings” which is a step of observing nucleic acid sequence data and separating each nucleic acid sequence string into smaller strings with a given length, which requires the mathematic concept of counting bases. Storing data is a routine step carried out by general-purpose computer systems. Associating a data structure with a sequence is a mental step of observing the sequence, determining the data structure should be associated, and annotating the index with the data structure and associated data. The “associated data structures” have no particular structural or functional aspects. MPEP 2106.05(a)(2))
10. (Previously Presented) The method of claim 1, wherein the first set of the plurality of nucleic acid sequences comprises one or more complete genomic sequences.
(EIA- data gathering limitation, describing an aspect of the data gathered.)
11. (Previously Presented) The method of claim 1, wherein the second set of the plurality of nucleic acid sequences comprises one or more complete genomic sequences.
(EIA- data gathering limitation, describing an aspect of the data gathered.)
12. (Currently Amended) A system for identifying a conserved target region in a subset of a plurality of nucleic acid sequences, the system comprising: one or more processors; memory storing one or more programs, the one or more programs configured to be executed by the one or more processors and including instructions to:
(EIA- routine general-purpose computer system)
The remainder of the analysis is the same as claim 1.
13. (Currently Amended) A non-transitory computer-readable storage medium storing one or more programs for identifying a conserved target region in a subset of a plurality of nucleic acid sequences, the one or more programs configured to be executed by one or more processors and including instructions to:
(EIA- routine general-purpose computer readable media)
The remainder of the analysis is the same as claim 1.
With respect to step 2A (2): NO, the claims do not integrate the JE into a practical application (MPEP 2106.04(d)):
“Examiners evaluate integration into a practical application by: (1) identifying whether there are any additional elements recited in the claim beyond the judicial exception(s); and (2) evaluating those additional elements individually and in combination to determine whether they integrate the exception into a practical application, using one or more of the considerations introduced in subsection I supra, and discussed in more detail in MPEP §§ 2106.04(d)(1), 2106.04(d)(2), 2106.05(a) through (c) and 2106.05(e) through (h).”
Claim(s) 1, 8, 10-13 recite the additional non-abstract element(s) of data gathering, or a description of the data gathered.
Data gathering steps are not an abstract idea, they are extra-solution activity, as they collect the data necessary to carry out the JE. MPEP 2106.05(g).
The data gathering does not impose any meaningful limitation on the JE, or how the JE is performed. MPEP 2106.05(g).
The data gathering steps constitute a general link to a technological environment: the trait prediction methods are intended to be applied to plant populations. (MPEP 2106.05(h), citing Mayo, Bilski, electric Power Group, Genetic Techs Ltd v Merial LLC.)
The additional limitation (data gathering) must have more than a nominal or insignificant relationship to the identified judicial exception to provide integration into a practical application. (MPEP 2106.05(g) citing Mayo, PerkinElmer, Inc. v. Interna Ltd, Intellectual Ventures LLC v. Erie Indem. Co., Electric Power Group LLC v. Alstom S.A.).
Claim(s) 1, 2, 9, 10-13 recite the additional non-abstract element (EIA) of a general-purpose computer system or parts thereof.
The claims do not provide any details of how specific structures of the computer elements are used to implement the JE. MPEP 2106.05(a), contrasting decisions identifying how the computer implements an abstract idea, such as in McRo to decisions which found no specific interaction with the computer, such as in Affinity Labs of Tex v. DirecTV, LLC.
The computer elements of the claims do not provide improvements to the functioning of the computer itself. MPEP 2106.05(a) I, contrasting decisions indicating an improvement to the computer, such as DDR Holdings, LLC v. Hotels.com LP, with decisions that did not identify an improvement to the computer, such as FairWarning IP, LLC v. Iatrix Sys.
The computer elements of the claims do not provide improvements to any other technology or technical field. MPEP 2106.05(a) II: contrasting decisions indicating an improvement to the technology, such as Diamond v. Diehr, Trading Techs. Int’l v. CQG Inc, or Intellectual Ventures I v. Symantec Corp, with decisions that did not identify an improvement to the technology, such as Alice Corp, Versata Dev. Group, Inc. v. SAP AM. Inc, or TLI Communications.
The computer elements of the claims do not utilize a particular machine. MPEP 2106.05(b): contrasting decisions wherein a particular machine was identified, such as MacKay Radio & Tel. Co. v. Radio Corp. of America, Eibel Process Co. v. Minn. & Ont. Paper Co., with decisions where a general-purpose computer does not qualify as a particular machine, such as Ultramercial, Inc. v. Hulu, LLC, TLI communications, or Eon Corp. IP holdings LLC v. AT&T Mobility LLC.
Hence, these are mere instructions to apply the JE using a computer, and therefore the claim does not recite integrate that JE into a practical application.
Dependent claim(s) 2-7, 9 recite(s) an abstract limitation to the JE reciting additional mathematic concepts, or mental processes. Additional abstract limitations cannot provide a practical application of the JE as they are a part of that JE.
In combination, the limitations of data gathering, for the purpose of carrying out the JE, using a general-purpose computer merely provide extra-solution activity, and fail to integrate the JE into a practical application.
With respect to step 2B: NO, the claims do not recite a specific inventive concept. The judicial exception alone cannot provide that inventive concept or practical application (MPEP 2106.05).
“… an "inventive concept" is furnished by an element or combination of elements that is recited in the claim in addition to (beyond) the judicial exception, and is sufficient to ensure that the claim, as a whole, amounts to significantly more than the judicial exception itself. Alice Corp…”
With respect to claim(s) 1, 8, 10-13: The limitation(s) identified above as non-abstract elements (EIA) related to data gathering do not rise to the level of significantly more than the judicial exception.
Glick (2009; PTO-1449) receives genomic sequence data comprising a plurality of nucleic acid sequences, including DNA.
Cardonha (2014; PTO-1449) receives genomic sequence data comprising a plurality of nucleic acid sequences, including DNA.
Ning (2001; PTO-1449) receives genomic sequence data comprising a plurality of nucleic acid sequences, including DNA.
Tsaltaris (2007; PTO-1449) receives genomic sequence data comprising a plurality of nucleic acid sequences, including DNA.
These elements meet the BRI of the identified data gathering limitations. As such, the prior art recognizes that this data gathering element is routine, well understood and conventional in the art. MPEP 2106.05(d): “If, however, the additional element (or combination of elements) is no more than well-understood, routine, conventional activities previously known to the industry, which is recited at a high level of generality, then this consideration does not favor eligibility.”
In the specification at [0061] it is disclosed that the steps identified as data gathering can be met by downloading or obtaining genomic sequence data from publicly available databases, such as those available at NCBI.
Data gathering steps are not an abstract idea, they are extra-solution activity, as they collect the data necessary to carry out the JE. MPEP 2106.05(g).
The data gathering does not impose any meaningful limitation on the JE, or how the JE is performed. MPEP 2106.05(g).
The additional limitation (data gathering) must have more than a nominal or insignificant relationship to the identified judicial exception to provide an inventive concept. (MPEP 2106.05(g) citing Mayo, PerkinElmer, Inc. v. Interna Ltd, Intellectual Ventures LLC v. Erie Indem. Co., Electric Power Group LLC v. Alstom S.A.)
The data gathering steps constitute a general link to a technological environment: the trait prediction methods are intended to be applied to plant populations. (MPEP 2106.05(h), citing Mayo, Bilski, electric Power Group, Genetic Techs Ltd v Merial LLC.)
Therefore, simply appending well-understood, routine, conventional activities previously known to the industry, specified at a high level of generality, to the judicial exception are insufficient to provide significantly more (as discussed in Alice Corp.,).
With respect to claim(s) 1, 2, 9, 12-13: the limitations identified above as non-abstract elements (EIA) related to general-purpose computer systems do not rise to the level of significantly more than the judicial exception.
Each of Glick, Cardonha, Ning and Tsaltaris disclose computer systems or computing elements which meet the BRI of the claimed computer system or computer system elements, comprising input, output/ display, a processor, and memory. These include computer system elements capable of storing 4 ^ k elements.
As such, the prior art recognizes that these computing elements are routine, well understood and conventional in the art.
The specification, at [0026-0036] discloses the use of routine general-purpose computers for carrying out the invention, and/or the use of commercially available computer system elements.
The claims do not provide any details of how specific structures of the computer elements are used to implement the JE. MPEP 2106.05(a), contrasting decisions identifying how the computer implements an abstract idea, such as in McRo to decisions which found no specific interaction with the computer, such as in Affinity Labs of Tex v. DirecTV, LLC.
The computer elements of the claims do not provide improvements to the functioning of the computer itself. MPEP 2106.05(a) I, contrasting decisions indicating an improvement to the computer, such as DDR Holdings, LLC v. Hotels.com LP, with decisions that did not identify an improvement to the computer, such as FairWarning IP, LLC v. Iatrix Sys.
The computer elements of the claims do not provide improvements to any other technology or technical field. MPEP 2106.05(a) II: contrasting decisions indicating an improvement to the technology, such as Diamond v. Diehr, Trading Techs. Int’l v. CQG Inc, or Intellectual Ventures I v. Symantec Corp, with decisions that did not identify an improvement to the technology, such as Alice Corp, Versata Dev. Group, Inc. v. SAP AM. Inc, or TLI Communications.
The computer elements of the claims do not utilize a particular machine. MPEP 2106.05(b): contrasting decisions wherein a particular machine was identified, such as MacKay Radio & Tel. Co. v. Radio Corp. of America, Eibel Process Co. v. Minn. & Ont. Paper Co., with decisions where a general-purpose computer does not qualify as a particular machine, such as Ultramercial, Inc. v. Hulu, LLC, TLI communications, or Eon Corp. IP holdings LLC v. AT&T Mobility LLC.
Hence, these are mere instructions to apply the JE using a computer, and therefore the claim does not provide significantly more.
Dependent claim(s) 2-7, 9 each recite a limitation requiring additional mathematic concepts or mental processes. Additional abstract limitations cannot provide significantly more than the JE as they are a part of that JE (MPEP 2106.05).
In combination, the data gathering steps providing the information required to be acted upon by the JE, performed in a generic computer or generic computing environment fail to rise to the level of significantly more than that JE. The data gathering steps provide the data for the JE, which is carried out by the general-purpose computers. No non-routine step or element has clearly been identified.
The claims have all been examined to identify the presence of one or more judicial exceptions. Each additional limitation in the claims has been addressed, alone and in combination, to determine whether the additional limitations integrate the judicial exception into a practical application. Each additional limitation in the claims has been addressed, alone and in combination, to determine whether those additional limitations provide an inventive concept which provides significantly more than those exceptions. For these reasons, the claims, when the limitations are considered individually and as a whole, are rejected under 35 USC § 101 as being directed to non-statutory subject matter.
Applicant’s Arguments:
Applicant’s arguments have been carefully considered but are not persuasive.
With respect to the arguments that the claims cannot be practiced in the human mind, these arguments are not persuasive. Of the limitations identified as mental steps, these limitations do not require any element for which the human mind is not equipped.
In the independent claims, the following limitations were determined to be mental processes:
“selecting a subset of the plurality of nucleic acid sequences, wherein the nucleic acid sequences in the subset share a characteristic;”
“identifying a target region that is conserved in the subset of nucleic acids, wherein the identifying comprises:
identifying, by the first index, a region of the target region as a conserved region appearing in every nucleic acid sequence in the first set;” and
determining, by the second index, that the conserved region appears in none of the nucleic acid sequences in the second set;”
These limitations require steps of observation, analysis and judgement, which fall within the grouping of Mental Process abstract ideas. MPEP 2106.04(a)(2).
With respect to the identification of certain steps as mental processes, and whether the human mind is equipped to perform the steps classified as mental processes: there is no step in the rejected claims which require an element which cannot practically be performed in the human mind because there are no limitations requiring an element such as a GPS receiver, specific computer-integrated elements such as network monitors or network packets, nor are there specific limitations to a multistep encryption of data for computer communication. No modified computer structures, such as self-referential tables, neural networks, artificial intelligence are present in the claims. (See SRI Int’l, Inc. v. Cisco Systems, Inc.,; CyberSource…)
MPEP 2106.04(a)(2) subsection 3: “… claims recite a mental process when they contain limitations that can practically be performed in the human mind, including for example, observations, evaluations, judgments, and opinions. Examples of claims that recite mental processes include:
• a claim to "collecting information, analyzing it, and displaying certain results of the collection and analysis," where the data analysis steps are recited at a high level of generality such that they could practically be performed in the human mind, Electric Power Group v. Alstom, S.A., 830 F.3d 1350, 1353-54, 119 USPQ2d 1739, 1741-42 (Fed. Cir. 2016);
• claims to "comparing BRCA sequences and determining the existence of alterations," where the claims cover any way of comparing BRCA sequences such that the comparison steps can practically be performed in the human mind, University of Utah Research Foundation v. Ambry Genetics, 774 F.3d 755, 763, 113 USPQ2d 1241, 1246 (Fed. Cir. 2014);…”
Applicant’s arguments in this aspect appear to be directed to an amount of information required to be processed. The amount of information generated at each step is not a sufficient process which takes the claim out of the realm of the human mind- the process may be long and tedious but there are no limitations requiring a particular manipulation of the data in any step that cannot be performed by mental processes as identified by the Courts.
The limitations of a) selecting sequences that share a characteristic, b) identifying target regions that are conserved by i) identifying and ii) determining a region that is conserved in one set and not another, do not require processes for which the human mind is not equipped. Identification of a shared characteristic and selection of a set of nucleic acid sequences that have that characteristic can be carried out in the human mind or using the computer as a tool, without any specialized computer elements required. Observing conserved or identical strings within one set can be carried out in the human mind or using the computer as a tool, without any specialized computer elements required. Determining the second set does not contain the conserved region can be carried out in the human mind or using the computer as a tool, without any specialized computer elements required. Each step observes, analyzes and makes a judgement on each piece of data, one at a time, to carry out the claimed process.
The amount of data, in and of itself is not a limitation which takes a process out of the realm of the human mind. It is the process performed on that data which is the mental step, and mental steps identified in the claims do not have to be fastest, most efficient, or require specialized computing elements. Data comparison, alignment, annotation, and output, can all be performed in the human mind, albeit very slowly, using pen and pencil, and slightly faster using the general-purpose computer as a tool or in a computing environment. Computations on large amounts of data performed mentally, or with paper and pencil, would take considerable time and effort, that is, of course, the singular purpose of computers and computer networks, to perform large numbers of calculations, via algorithms, rapidly, and without error (assuming no error in user input). Although a general-purpose computer can perform calculations at a rate and accuracy that can far outstrip the mental performance of a skilled artisan, the nature of the activity is essentially the same, and constitutes an abstract idea. See Bancorp Serves., L.L. C. v. Sun Life Assur. Co. of Canada (U.S.) (holding that “the fact that the required calculations could be performed more efficiently via a computer does not materially alter the patent eligibility of the claimed subject matter”); see also SiRF Tech., Inc. v. Int’l Trade Comm ’n, (Fed. Cir. 2010) (holding that: In order for the addition of a machine to impose a meaningful limit on the scope of a claim, it must play a significant part in permitting the claimed method to be performed, rather than function solely as an obvious mechanism for permitting a solution to be achieved more quickly, i.e., through the utilization of a computer for performing calculations)."
The creation of the two indices has been identified as a mathematic concept and not a mental process. The creation of the indexes utilizes mathematic relationships, mathematic formulae and mathematic calculations. (MPEP 2106.04(a)(2) section 1. Examples similar to the identified mathematic concepts: Benson; Mackay Radio & Tel. Co. v. Radio Corp. of America; Bilski v. Kappos, ; SAP America, Inc. v. InvestPic, LLC, ; Parker v. Flook; Burnett v. Panasonic Corp., etc.)
MPEP 2106.04(a)(2): “It is important to note that a mathematical concept need not be expressed in mathematical symbols, because "[w]ords used in a claim operating on data to solve a problem can serve the same purpose as a formula." In re Grams, 888 F.2d 835, 837 and n.1, 12 USPQ2d 1824, 1826 and n.1 (Fed. Cir. 1989). See, e.g., SAP America, Inc. v. InvestPic, LLC, 898 F.3d 1161, 1163, 127 USPQ2d 1597, 1599 (Fed. Cir. 2018) (holding that claims to a ‘‘series of mathematical calculations based on selected information’’ are directed to abstract ideas); Digitech Image Techs., LLC v. Elecs. for Imaging, Inc., 758 F.3d 1344, 1350, 111 USPQ2d 1717, 1721 (Fed. Cir. 2014) (holding that claims to a ‘‘process of organizing information through mathematical correlations’’ are directed to an abstract idea); and Bancorp Servs., LLC v. Sun Life Assurance Co. of Can. (U.S.), 687 F.3d 1266, 1280, 103 USPQ2d 1425, 1434 (Fed. Cir. 2012) (identifying the concept of ‘‘managing a stable value protected life insurance policy by performing calculations and manipulating the results’’ as an abstract idea).”
Creating permutations of a nucleic acid sequence utilizes combinatorial mathematics. Calculating permutation index values for k-mers extracted from a nucleic acid sequence generates a unique integer that represents the specific sequence’s position within an ordered list of all possible k-mers of that length. Each nucleotide is assigned a numerical value based on a given order. As nucleic acids use a 4 letter alphabet (ACTG), calculating a specific k-mer’s permutation index value converts a base-4 (quaternary) number into a standard base-10 integer. For a k-mer of length k, the total number of unique permutations is 4^k. The permutation indices will range from 0 to 4^k - 1. The index is calculated by summing the value of each nucleotide multiplied by its positional weight 4^position. As such, the limitations to creating the indices which contain the permutations each recite a mathematic concept. MPEP 2106.04(a)(2).
The indexes themselves and associated data structures were identified as possible field of use limitations, because the indexes have no particular structure, function or activity, and the associated data structures have no particular structure, function or activities. In the independent claims, the first and second indices have the following identical descriptions:
“a first index representing a first set of the plurality of nucleic acid sequences corresponding to the subset, wherein the first index comprises at least 412 elements representing each respective permutation of nucleic acid sequences, and wherein the data created and stored in the first index comprises a first plurality of data structures each associated with a respective nucleic acid sequence of the first set;”
The index has no particular data structure constructed to store the 412 pieces of data in any particular way. The index does not create the permutations of each nucleic acid sequence. The index does not create the associated data structures. The associated data structures do not have any particular structural elements, or particular information to be stored in association with a nucleic acid sequence or permutated nucleic acid sequence. The associated data structures appear to merely store “associated data” with no further action performed within, or by the associated data structures. Both types of “structures” can be met by flat databases, tables or spreadsheets.
Applicant is pointed to the prior art for examples of index data structures for k-mer indexing which provide structural and functional details about the index. (For example, Kahvecki et al. (2001) An Efficient Index Structure for String Databases. Proceedings of the 27th VLDB Conference, Roma Italy, 10 pages. See Section 3.3/ fig 5. Hunt et al. (2001) A Database Index to Large Biological Sequences. Proceedings of the 27th VLDB Conference, Roma Italy, 10 pages (Previously provided, PTO-892). Ozturk et al. (2003) Effective Indexing and Filtering for Similarity Search in Large Biosequence Databases. Third IEEE symposium on Bioinformatics and Bioengineering, 12 March 2003, Bethesda MD, USA, 8 pages. Lopresti et al. (2013) Permutation Index and GPU to Solve Many Queries. HPCLatAm, Session: GPU architecture and applications, C Garcia Garino and M Printista (Eds) Mendoza Aggentina, July 29-30, 2013, p101-112. See section 4, Fig 1. (Previously provided, PTO-892))
The limitations identified as elements in addition to the judicial exception are steps of data gathering, and the use of general-purpose computers. Both were shown above to be insufficient to integrate the judicial exception into a practical application, and both were shown above to be insufficient to provide a specific inventive concept.
Further, with respect to the arguments regarding the alleged improvement, it is unclear that the independent claims recite all the necessary and sufficient steps required to achieve that improvement. MPEP 2106.05(a): “An important consideration in determining whether a claim improves technology is the extent to which the claim covers a particular solution to a problem or a particular way to achieve a desired outcome, as opposed to merely claiming the idea of a solution or outcome. McRO, 837 F.3d at 1314-15, 120 USPQ2d at 1102- 03; DDR Holdings, 773F.3d at 1259, 113 USPQ2d at 1107.”
The MPEP sets forth that “if the examiner concludes the disclosed invention does not improve technology, the burden shifts to applicant to provide persuasive arguments supported by any necessary evidence to demonstrate that one of ordinary skill in the art would understand that the disclosed invention improves technology. Any such evidence submitted under 37 CFR 1.132 must establish what the specification would convey to one of ordinary skill in the art and cannot be used to supplement the specification.” Applicant’s arguments cannot take the place of evidence.
Claim Rejections - 35 USC § 112
The following is a quotation of 35 U.S.C. 112(b):
(b) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention.
The following is a quotation of 35 U.S.C. 112 (pre-AIA ), second paragraph:
The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the applicant regards as his invention.
Claims 1-13 remain rejected under 35 U.S.C. 112(b) or 35 U.S.C. 112 (pre-AIA ), second paragraph, as being indefinite for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor (or for applications subject to pre-AIA 35 U.S.C. 112, the applicant), regards as the invention.
The metes and bounds of claims 1, 12 and 13 are unclear with respect to the data elements stored in each index. The indices are formed, as simply a first subset of a plurality of nucleic acid sequences, and “respective” permutations, and a second subset of a plurality of nucleic acid sequences, and “respective” permutations. The subsets have no further description. The subsets are not chosen on any particular basis. The newly added limitation to “sharing a characteristic” fails to set forth what characteristic is intended to be shared within the first subset. Nucleic acid sequences can have a multiplicity of characteristics, including sequence, length, type of nucleic acid (DNA, RNA, mRNA etc.), coding sequence, molecular weight, location in a genome, presence or absence of mutations, secondary and tertiary structure characteristics, presence or absence of methylation, presence or absence of a motif, etc. (Wikipedia, 2026).
The two indices have no intended commonality or differences, beyond the “first” and “second” subsets and the newly added limitation that the second set is “distinct from the first (sub)set.” The metes and bounds of “distinct” here are unclear as it is not clearly that none of the sequences or permutations of the first set are in the second set- merely that the two sets are different from one another in some manner.
A review of Fig 3A indicates that certain necessary steps may be missing from claim 1, as claim 1 does not recite the requirements of element 306 (extracting substrings of length k, where k is the number of bases), nor the requirements of element 308 (storing data representing each extracted substring wherein “the reference data associates position data of the substring, identity of the nucleic acid sequence and an element of the first index with one another), nor the requirements of element 312, (comparison of “data stored in the first index with the same element and for a corresponding position” for every other sequence in the subset).
Further, if every respective possible permutation of a sequence is generated in the first index, there will be no shared region that is “conserved (e.g. matching, identical)” in every sequence of the first (sub)set. The specification defines respective permutations as:
“[0044] … In the case of FIG. 2, each permutation is 16 bases in length, resulting in an index with 416 or 4,294,967,296 elements (note that each base of a nucleic acid sequence is one of four types). More generally, the size or the number of elements of index 200 is equal to 4k, where k is the length, in bases, of each permutation.”
Conserved is defined as:
“[0023] … The indexes may be used to quickly, efficiently, and accurately locate all regions of a predefined minimum length that are both conserved (e.g., matching, identical) across all of the nucleic acid sequences in the first index and unique against all nucleic acid sequences in the second index.”
As shown in this toy example below, there is no conserved, matching or identical region that is present in all nucleic acid sequences in the first set: k = 3, 4^3 = 64, 63 permutations of GGG to reach CCC. Each 3-mer is different.
PNG
media_image1.png
962
598
media_image1.png
Greyscale
A review of the Figures and specification suggests that this is intended to select a subset of the nucleic acid sequences in the first index, however the claim does not specify this interpretation. See Fig 3A and 3B. Fig 3A states that the “identifying” step is to “locate conserved regions of length l that appear in all nucleic acids of the subset” in elements 310 and 316, which are not reflected in the independent claims.
For example, in the subset of the first 16 entries in my toy index set above, four entries share the conserved region of length 2 nucleotides (l = 2), where “GG” are the conserved nucleotides and l < k.
PNG
media_image2.png
332
194
media_image2.png
Greyscale
Additionally, Fig 3B suggests that the second index is created using all the sequences not selected for the first subset at element 322 and element 324. The independent claims merely indicate the two subsets are distinct. The generation of the second index is further lacking the specific requirements of elements 322 and 324, before moving on to element 326. The remaining nucleic acid sequences not selected for the first index all have kmers extracted of length k (element 322), and the second index comprises specific data (element 324) “the reference data associates position data of the substring, identity of the nucleic acid sequence and an element of the second index with one another.” Figure 3B discusses steps required for the iterative processing which are lacking from the independent claims (elements 328, 332, 336, 338). While claims are read in light of the specification, limitations from the specification cannot be read into the claims.
The metes and bounds of claims 1, 12 and 13 are unclear with respect to the term “respective permutations” of each nucleotide sequence in the set of nucleotide sequences used for each index. This does not clearly point out and distinctly claim what particular permutations are to be generated for each received sequence from each subset of the “plurality of nucleic acid sequences” received in the first step. Using “respective” fails to provide a type or description of the permutations. It is unclear if this is intended to represent known SNP or SNV data, or other known variant information, or whether this step generates every possible permutation of a nucleic acid sequence where each nucleotide, in turn, is changed to another, across the entire set of sub-strings. This second interpretation appears to be the case, as each created index must comprise at least 412 elements/ permutations related to the initial set of sequences, but the claim does not clearly reflect that interpretation. The sequences used to create the first index are not all clearly from the same region of a genome, the same chromosome, or even the same genome. The sequences used to create the second index are not all clearly from the same region of a genome, the same chromosome, or even the same genome. The “identifying” does not clearly require any permutation data, associated data structures, nor does it set forth and distinctly claim any particular type or way to make the identification.
The metes and bounds of claims 1- 13, overall are unclear, with respect to how the created indices affect the remainder of the steps of the claims. The purpose of creating the two indices with all the permutations is entirely unclear with respect to the other steps performed by the independent claims. The 412 elements and associated data structures created for each index are not clearly used in any positive active method step of the independent claims, nor does their presence provide any particular information, data, or data structure necessary to carry out any subsequent step. The “identifying by the first index” does not clearly require any permutation data or the associated data structures, nor does it set forth and distinctly claim any particular type or way to make the identification. The “determining” step allegedly uses “by the second index” however it is entirely unclear how this is to be performed, and it is further unclear how this would provide any relevant information, as the point of the method appears to be to identify a common region in one subset of nucleic acid sequences, that is not present in another subset of nucleic acid sequences: this is not the comparison of the identified common region with the permutations in the second index, as the method is not attempting to identify or confirm the presence or absence of the permutated sequences, but the presence or absence of the common region with the second set of nucleic acid sequences from the initial plurality. The point of creating these indices is entirely unclear.
The metes and bounds of claims 2 and 9 are unclear with respect to how the sequences are divided and permutated. Claim 2 sets forth that the creation of the first index comprises “dividing the (each) nucleic acid sequence into a plurality of sub-strings” without any further description as to how or why any sequence would be subdivided into sub-strings. Claim 9 sets forth that the creation of the second index comprises “dividing the (each) nucleic acid sequence into a plurality of sub-strings” without any further description as to how or why any sequence would be subdivided into sub-strings. The partitioning of each nucleic acid sequence appears to be random, and can encompass from substrings of a single nucleotide, up to the full length of the genome. The specification suggests this is the generation of k-mers of a given length, however this is not clearly recited (see Fig 3). Further, in claims 2/9, each substring “stores” data structures which comprise information not clearly provided by claim 1, from which claims 2/9 depend. The data stored in the data structure is intended to be a) an identity of the original nucleic acid from which each substring was created, b) “a permutation of bases forming the substring” and c) a position of the substring in the original nucleic acid sequence. It is entirely unclear how these elements are determined or generated with the data at hand. The data at hand is merely “genomic data” which comprised “a plurality of nucleic acid sequences.” No additional information was clearly provided. Claims 2 and 9 do not set forth how any of the information, such as the identity of the nucleic acid sequence, is obtained or determined. Further, the metes and bounds of “a permutation of bases forming the substring” are unclear. It is unclear if this is intended to be the identification of known variant data (i.e. SNP, SNV), from some unknown, outside source, or whether this is intended to record a change made in the creation of the 412 members of the index. This step fails to set forth how the permutations were made, or what the permutations are intended to be. It is further unclear how any of these limitations further affect any step of claim 1, as the claims fail to set forth and particularly claim how the permutations, the data structures, or any other information are used to carry out the “identifying, by the first index” “determining by the second index” and “generating and outputting” steps of claim 1.
The metes and bounds of claim 3 are unclear, with respect to how identifying a common (unspecified) permutation between two sub-strings of a first nucleic acid from the first set leads to “identifying the region appearing in every nucleic acid in the first set” as required. Only two sub-strings are analyzed from the first subset of sequences, not all sub-strings. If these two sequences have a difference (permutation) from the “genomic data comprising a plurality of nucleic acid sequences” then there is no conserved region. The claim does not set forth what the permutation is, or how it is identified, or how that identification leads to the identification of a conserved region in the first subset. It is unclear if this is a matching of the “associated data structure” information and data values, or whether this is intended to be a matching operation on the sequence string data of the sub-strings, or some other procedure. It is not clearly a mapping process where a genomic location is identified for the sub-string, where the presence of a mutation in that sub-string might be compared to known SNV data for that genomic location. There is no expectation that any two nucleic acid sequences from that first subset of nucleic acid sequences would be from the same genomic region, or that they would overlap, or have a common permutation in either permutation index.
The metes and bounds of claims 4-5 are unclear with respect to how the second data structure from claim 3, which was associated with the second nucleic acid sequence from the first subset of nucleic acid sequences, could have “previously” been matched to a sub-string. No data representing prior matching operations is clearly provided by any other claim. How any time-dependent parameter is to be determined is completely unclear. The limitations to “at a span occurring immediately before the given sub-string” are completely unclear, as there was no expectation that the second nucleic acid sequence from that first subset of nucleic acid sequences would be from the same genomic region, or that they would overlap, or have some positional relationship with the first nucleic acid sequence, or have a common permutation in the permutation index. What the sub-strings are intended to “span” is unclear. The sub-strings are not made in any particular way, nor in any particular order. Claim 5 allegedly sets forth that they are adjacent to one another within the first nucleic acid sequence, without providing how this is determined with the data at hand. Claim 5 appears to compare elements of the first index with a different set of substrings of the first nucleic acid sequence. How this occurs is unclear. It is further unclear how the second data structure comprises any of the required information related to the previously-matched strings, and any positional information, as these “associated data structures” are generated in claim 1 without any of that knowledge.
The metes and bounds of claim 6 are unclear with respect to how the determination of a discriminatory region is actually carried out, when the steps only compare data within the first set (a common permutation in the first set, an identity in the first set with positional information, and a position from the second set that locates the substring “immediately after” the previously matched substring of the first set). If the determination finds that substring A meets these conditions, it is entirely unclear how it is determined not to be within the second subset of nucleic acid sequences.
The metes and bounds of claim 7 are unclear with respect to the iteration, the arrangement of the (random) sub-strings, and “failing” one of the “three conditions” from claim 6. It is entirely unclear how the elements of the claim “determine that, for every nucleic acid sequence in the second index, at least one data structure fails at least one of the three conditions…” Claim 6 does not clearly set forth a process capable of being iterated to achieve this goal. Claim 6 sets forth conditions, but not how they are to be compared, or used to analyze the 412 substrings and 412 associated data structures of the indexes in any iterative manner.
The metes and bounds of claims 10-11 are unclear with respect to how the description of the first or second set of nucleic acid sequences as “comprises one or more complete genomic sequences.” The metes and bounds of the term “complete” are unclear with respect to how it differs from the genomic sequence information already obtained. It is unclear if these are multiple copies of the same complete genomic sequence (repeated copies of HG26), different genomes of the same species (individual complete genomic sequences from more than one human), or whether this is intended to encompass different genomes of different species/ organisms (mouse and man). It is further unclear how the presence of a full genome affects the generation of the indexes of claim 1, or the identification of a “region” present in one index.
Applicant’s Arguments:
Applicant’s arguments and amendments related to this rejection have been considered but are not persuasive.
The MPEP sets forth two requirements under 35 USC 112b at MPEP 2171:
“(A) the claims must set forth the subject matter that the inventor or a joint inventor regards as the invention; and
(B) the claims must particularly point out and distinctly define the metes and bounds of the subject matter to be protected by the patent grant.”
“The first requirement is a subjective one because it is dependent on what the inventor or a joint inventor for a patent regards as his or her invention. ...”
“The second requirement is an objective one because it is not dependent on the views of the inventor or any particular individual, but is evaluated in the context of whether the claim is definite — i.e., whether the scope of the claim is clear to a hypothetical person possessing the ordinary level of skill in the pertinent art.”
The examiner has reviewed the broadest reasonable interpretation of the amended claims, in light of the specification, and identified specific issues remaining in the claims (MPEP 2173.02). The examiner has pointed out multiple inconsistent interpretations for the limitations of the claims and provided reasoning the claims remain indefinite (MPEP 2173.02). The claims fail to provide the necessary and sufficient limitations to carry out the desired steps and achieve the desired results.
MPEP 2173.02: “The essential inquiry pertaining to this requirement is whether the claims set out and circumscribe a particular subject matter with a reasonable degree of clarity and particularity. "As the statutory language of ‘particular[ity]' and 'distinct[ness]' indicates, claims are required to be cast in clear—as opposed to ambiguous, vague, indefinite—terms. It is the claims that notify the public of what is within the protections of the patent, and what is not." Packard, 751 F.3d at 1313, 110 USPQ2d at 1788.
Further, Applicant’s arguments are directed to limitations not present in the rejected claims.
MPEP 2111: "Though understanding the claim language may be aided by explanations contained in the written description, it is important not to import into a claim limitations that are not part of the claim. For example, a particular embodiment appearing in the written description may not be read into a claim when the claim language is broader than the embodiment." Superguide Corp. v. DirecTV Enterprises, Inc., 358 F.3d 870, 875, 69 USPQ2d 1865, 1868 (Fed. Cir. 2004).
Claim Rejections - 35 USC § 102
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action:
A person shall be entitled to a patent unless –
(a)(1) the claimed invention was patented, described in a printed publication, or in public use, on sale, or otherwise available to the public before the effective filing date of the claimed invention.
(a)(2) the claimed invention was described in a patent issued under section 151, or in an application for patent published or deemed published under section 122(b), in which the patent or application, as the case may be, names another inventor and was effectively filed before the effective filing date of the claimed invention.
Claim(s) 1- 13 remain rejected under 35 U.S.C. 102a1 as being anticipated by Glick et al (2009).
Glick, B. et al. Method for Indexing Nucleic Acid Sequences for Computer Based Searching. US 2009/0270277 A1, published 10/29/2007.
With respect to claims 1, 12 and 13, Glick meets: “Claim 1. A method for identifying a conserved target region in a subset of a plurality of nucleic acid sequences, comprising: at a system comprising one or more processors and memory storing instructions executable by the processor: receiving genomic data representing a plurality of nucleic acid sequences;”
Glick receives or obtains genomic sequences, such as the complete human genome, at [0006, 0028, 0048, 0051, 0055], from a database such as Ensembl. The nucleic acid sequences can be DNA (both strands), RNA, mRNA etc [0024]. Glick provides computer implemented methods, computer systems, and non-transitory computer readable media throughout, see at least: [0005-0010, 0020-0021]. Glick also provides data structures stored on the computers, see at least: [0011, 0022].
With respect to claims 1, 12 and 13, Glick meets: “selecting a subset of the plurality of nucleic acid sequences, wherein the nucleic acid sequences in the subset share a characteristic;”
Glick selects sequences from the obtained genomic sequences, such as nucleic acid sequences of human chromosome 1, from the complete human genome, at least at [0025, 0028-0029] where shared characteristics can include: linear or circular topology [0027], length [throughout], position [0009, 0026], strandedness [0031]. Additional characteristics are set forth at [0030].
“[0030] … A MICA file consists of a Sequence Segment (containing values for the following fields: Segment Format, Segment Size, Sequence Properties, and the DNA Sequence (elements A-D of FIG. 1)) followed by an Index Segment (containing values for the following fields: Index Segment Format, Index Segment Size, Index Properties, Chunk Counts Summary, degenerate K-mer Count, N-Stretch Count (S), Chunk Data Array, Degenerate Data Array, and N-Stretch Data Array (elements E-M in FIG. 1)).”
With respect to claims 1, 12 and 13, Glick meets: “creating and storing data in a first index representing a first set of the plurality of nucleic acid sequences corresponding to the subset, wherein the first index comprises at least 412 elements representing each respective permutation of nucleic acid sequences, and wherein the data created and stored in the first index comprises a first plurality of data structures each associated with a respective nucleic acid sequence of the first set; creating and storing data in a second index representing a second set of the plurality of nucleic acid sequences distinct from the subset, wherein the second index comprises at least 412 elements representing each respective permutation of nucleic acid sequences, and wherein the data created and stored in the second index comprises a second plurality of data structures each associated with a respective nucleic acid sequence of the second set; and”
Glick creates at least two indices (Fig 1), called MICA indexes, for selected distinct subsets of nucleic acids from the initially received plurality. Glick discloses selecting a subset of polynucleotide sequences from the genomic sequence data at [0007, 0028-0029, 0051-0053] and creating indexes from that subset [0030-0036]. The genome can be broken into subsets (by chromosome) or chunks, representing separate distinct subsets of the genomic nucleic acid sequences [0008, 0055]. For each chunk or subset, Glick generates all degenerate substrings of the selected subset of nucleic acid sequences of a particular length K (K-mer, or Kmer) which meets the BRI of a “substring” or “sub-string”. The method then determines all of the fully or partially degenerate permutations of the base sequences of length K [0037 Query length Q), 0056 (3-mers, 4-mers, 6-mers, 8-mers), 0060-0061 (Query length), 0064-0065]. The Kmers may have degenerate or permutated bases [0007, 0025-0042]. The computer stores each Kmer sequence, each permutation of each Kmer sequence, in an associated data structure. Glick additionally stores pointers, and the position (relative or absolute) of each Kmer in the overall nucleotide sequence of the genome [0009, 0026].
[0009] “The position of each non-degenerate K-mer is recorded in a non-degenerate data array in the file. In one embodiment, the non-degenerate data array is divided into 4K partitions corresponding to all of the possible 4K non-degenerate K-mers. Each partition contains a list of integers representing the number of times a particular non-degenerate K-mer is present in each of the chunk sections, followed by a list of integers representing intra-chunk section positions of the particular K-mer in each of the chunk sections. The position of each partially degenerate K-mer is recorded in a degenerate data array in the file. Each particular partially degenerate K-mer is represented as an integer that marks the absolute position of the particular K-mer, followed by a string that encodes the sequence of the particular K-mer.”
Each index for each chunk of the genomic sequence data received by Glick meets the BRI of the two created indexes of claims 1, 12 and 13. The indexes of Glick can comprise 4k elements. The indexes of Glick can comprise entire genomic sequences, as shown in Fig 2, and [0006], or subsets of genomic sequences, such as only Chromosome 1. The kmers of Glick can comprise permutations (degenerate bases) [0005]. The indexes can comprise additional data structures with information linked to the sequence.
“[0011] In another embodiment, the invention provides a data structure for recording information regarding a nucleic acid, the information being stored in a computer file. The data structure contains information on the K-mer's of the nucleotide sequence. The data structure comprises a non-degenerate data array containing the position of each particular non-degenerate K-mer of a nucleotide sequence; a degenerate data array containing the position of each particular partially degenerate K-mer and a string encoding the sequence of each particular partially degenerate K-mer; a sequence segment format field that contains an integer identifying this segment as the sequence segment; a sequence segment size field containing an integer representing the total number of bytes occupied by the sequence segment; a sequence properties field containing an integer representing the topology and the strandedness of the nucleotide sequence; a DNA sequence field containing the base sequence of the nucleotide sequence; an index segment format field that contains an integer identifying this segment as the index segment; an index segment size field containing an integer representing the total number of bytes occupied by the index segment; an index properties field containing an integer representing the byte order of the index; a chunk counts summary field containing a list of integers representing the total number of times a particular non-degenerate K-mer appears in the nucleotide sequence; a degenerate K-mer count field containing an integer representing the total number of partially degenerate K-mers in the nucleotide sequence; an N-stretch count field containing an integer S representing the number of separate stretches of K or more consecutive N's in the nucleotide sequence; and an N-stretch data array containing S pairs of integers that represent the starting positions and lengths of the separate stretches of consecutive N's in the nucleotide sequence.”
“[0029] The main body of a MICA index is the Chunk Data Array, which stores the positions of the nondegenerate K-mers (FIG. 1). The total number of position values is largely independent of K. However, there are 4.sup.K different nondegenerate K-mers, so the Chunk Data Array is divided into 4.sup.K partitions. Each partition is divided into C sub-partitions that contain the intra-chunk position values. The sizes of these sub-partitions are recorded in a list at the beginning of the partition.”
Glick meets all the requirements for the two indexes as set forth above.
With respect to claims 1, 12 and 13, Glick meets: “identifying a target region that is conserved in the subset of nucleic acids, wherein the identifying comprises: identifying, by the first index, a region of the target region as a conserved region appearing in every nucleic acid sequence in the first set; determining, by the second index, that the conserved region appears in none of the nucleic acid sequences in the second set; and
Glick provides identifying a Kmer of a given length, that appears in a chunk of the genomic nucleic acid sequence data, and how many times it appears [0012].
“[0012] In another embodiment of the invention, the invention provides a method of searching for a specific base sequence in a nucleotide sequence using a computer. In one embodiment, the invention comprises using a computer to access the data structure of the invention containing information on a nucleotide sequence and having the computer search the information in the data structure for the presence and location of the specific base sequence. The computer accesses a data structure file of the present invention containing information on the nucleotide sequence desired. The computer then loads the data from the index segment format field, the index segment size field, the index properties field, the chunk counts summary field, the degenerate K-mer count field and the N-stretch Count field into main memory. The computer then divides the specific base sequence that is being searched for into specific K-mers. The computer accesses the data loaded into the main memory and the information in the degenerate data array, the non-degenerate data array, and the N-stretch data array and uses this information to generate the position of the unique K-mers in the nucleotide sequence. The computer then uses this K-mer position data to determine the location of the specific base sequence in the nucleotide sequence if present.”
“[0025] … The absolute positions of a K-mer within a full DNA sequence can then be calculated with the aid of a list specifying the number of instances of the K-mer within each chunk.”
“[0037] The query length Q can range from one base to the length of the subject DNA sequence. Both strands of the DNA molecule are searched. For a query that is palindromic--i.e., identical to its reverse complement--a single search is performed. For a query that is nonpalindromic, two successive searches are performed, one with the query and another with the reverse complement of the query. If the DNA molecule is circular, the initial search is followed by a secondary search for matches that span the origin. One step in this secondary search involves dividing the query in half and then checking for one of two possibilities: either the first half-query matches within the last Q-1 bases of the DNA sequence, or the second half-query matches within the first Q-1 bases of the DNA sequence.”
Glick identifies sequences which exist in a first index, and not in the second index. Glick notes that:
[0039] “The index can be searched to return a complete list of exact matches for a nondegenerate or partially degenerate query of any length.” [0006, 0037-0043]
[0039] “… This strategy of starting with the rarest K-mer can significantly accelerate searches because some K-mers are found less frequently than others and therefore result in fewer comparisons. In chromosome 1, the most common 4-mer (AAAA) appears 56 times more often than the rarest 4-mer (CGCG), and the most common 6-mer (TTTTTT) appears 929 times more often than the rarest 6-mer (CGTACG).”
[0041] “… A search for GDGCHC will therefore return matches for all of the 49 possible matching 6-mers. Alternatively, searches for a partially degenerate query can be restricted to return only literal matches to that character string, so a search for "GDGCHC" will return matches only for the single 6-mer GDGCHC.”
[0043] “… MICA therefore uses an alternative intersection algorithm for partially degenerate K-mers. A boolean array of 65,535 elements is used to represent the positions in a chunk. For a given chunk, all of the individual K-mer lists are scanned, and the 2-byte position values are recorded by setting the corresponding boolean elements to true, yielding a boolean array that indicates which positions in the chunk match one of the K-mers. Then the intersection is obtained by checking whether each working list element corresponds to a value of true in the boolean array. This method is efficient due to the relatively small number of operations and the sequential nature of the memory accesses.”
Glick meets the BRI of the identifying steps as set forth above.
With respect to claims 1, 12 and 13, Glick meets: “generating and outputting data representing the identified conserved target region.”
Glick stores and outputs the desired data, using the general-purpose computer system, throughout. (General purpose computer system, and display [0012, 0019-0023, 0048] and the claims such as claim 34).
With respect to claims 2 and 9, and the creation of the indexes, Glick teaches that for each nucleic acid sequence in the chunk, Kmers are generated which meet the BRI of sub-strings, as set forth above. For each sub-string, Glick stores information regarding the identity of the sequence [0071], a permutation as set forth above, and a position, as set forth above.
With respect to claim 3, Glick teaches that common permutations can be identified within substrings.
With respect to claim 4, Glick teaches that identities can be determined by comparison or previous matching operations at [0071]. Location of the substring in the overall sequence can be determined as set forth above.
With respect to claims 5-7, Glick iterates their method with different lengths of Kmers, which are different substrings, to identify sequences containing a certain number of Kmers in a certain order.
With respect to claim 8, Glick indexes any kind of genetic sequence information, including at least DNA.
Claims 10-11 are met by the indexing of the complete human genome as set forth above.
Applicant’s arguments:
Applicant’s arguments have been carefully considered but are not persuasive. As set forth above, Glick meets the broadest reasonable interpretation of the claims, including identifying matching or conserved regions, and the creation of distinct indices.
Conclusion
THIS ACTION IS MADE FINAL. Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a).
A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to MARY K ZEMAN whose telephone number is 5712720723. The examiner can normally be reached on 8am-2pm M-F. Email may be sent to mary.zeman@uspto.gov if the appropriate permissions have been filed.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Larry Riggs can be reached on 571 270-3062. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of an application may be obtained from the Patent Application Information Retrieval (PAIR) system. Status information for published applications may be obtained from either Private PAIR or Public PAIR. Status information for unpublished applications is available through Private PAIR only. For more information about the PAIR system, see http://pair-direct.uspto.gov. Should you have questions on access to the Private PAIR system, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative or access to the automated information system, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/MARY K ZEMAN/ Primary Examiner, Art Unit 1686