DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Drawings
The drawings are objected to under 37 CFR 1.83(a) because they fail to show label 202, 211, 213, 215-219, 221-223, 225-227, 229, 231-234, 311, 313, 315-327, 329-332, 602, 613 as described in the specification. Any structural detail that is essential for a proper understanding of the disclosed invention should be shown in the drawing. MPEP § 608.02(d). Corrected drawing sheets in compliance with 37 CFR 1.121(d) are required in reply to the Office action to avoid abandonment of the application. Any amended replacement drawing sheet should include all of the figures appearing on the immediate prior version of the sheet, even if only one figure is being amended. The figure or figure number of an amended drawing should not be labeled as “amended.” If a drawing figure is to be canceled, the appropriate figure must be removed from the replacement sheet, and where necessary, the remaining figures must be renumbered and appropriate changes made to the brief description of the several views of the drawings for consistency. Additional replacement sheets may be necessary to show the renumbering of the remaining figures. Each drawing sheet submitted after the filing date of an application must be labeled in the top margin as either “Replacement Sheet” or “New Sheet” pursuant to 37 CFR 1.121(d). If the changes are not accepted by the examiner, the applicant will be notified and informed of any required corrective action in the next Office action. The objection to the drawings will not be held in abeyance.
The drawings are objected to as failing to comply with 37 CFR 1.84(p)(4) because reference character “212” has been used to designate both “Document Intelligence: A way to improve efficiency and productivity” and "Applicable Fields". Corrected drawing sheets in compliance with 37 CFR 1.121(d) are required in reply to the Office action to avoid abandonment of the application. Any amended replacement drawing sheet should include all of the figures appearing on the immediate prior version of the sheet, even if only one figure is being amended. Each drawing sheet submitted after the filing date of an application must be labeled in the top margin as either “Replacement Sheet” or “New Sheet” pursuant to 37 CFR 1.121(d). If the changes are not accepted by the examiner, the applicant will be notified and informed of any required corrective action in the next Office action. The objection to the drawings will not be held in abeyance.
Specification
The disclosure is objected to because of the following informalities: Figure 1B introduces the label 150 but has not been disclosed in the specifications. Appropriate correction is required.
The disclosure is objected to because of the following informalities: Figure 2 indicates the label of 218 as the "the first two bullet points of the body of the text, e.g., segments 216 and 218" and label 212 appears to claim two points on the figure as the title, what seems to be a bullet point. The Label 212 should be indicated 216 as demonstrated in specifications and the drawings should reflect this. Appropriate correction is required.
The disclosure is objected to because of the following informalities: Figure 4 introduced the label 400 but has not been disclosed in the specification. Appropriate correction is required.
The disclosure is objected to because of the following informalities: Figure 11A introduced the label 1100A but has not been disclosed in the specification. Appropriate correction is required.
The disclosure is objected to because of the following informalities: Figure 11B introduced the label 1100B but has not been disclosed in the specification. Appropriate correction is required.
Claim Rejections - 35 USC § 112
The following is a quotation of 35 U.S.C. 112(b):
(b) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention.
The following is a quotation of 35 U.S.C. 112 (pre-AIA ), second paragraph:
The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the applicant regards as his invention.
Claim(s) 8-10, 13, and 22 rejected under 35 U.S.C. 112(b) or 35 U.S.C. 112 (pre-AIA ), second paragraph, as being indefinite for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor (or for applications subject to pre-AIA 35 U.S.C. 112, the applicant), regards as the invention.
Claim 8 recites the limitation "wherein the user interface" and depends on claim 1 in which case, does not introduce “the user interface”. Claim 7 does introduce the use of “a user interface”. There is insufficient antecedent basis for this limitation. The claim first introduces “user interface” using the definitive article “wherein the user interface” without having previously introduced the element with an indefinite article “a user interface” earlier in the claim.
Claim 9-10 recites the limitation "the eye-tracking data" and nowhere in the claims is there a condition that recites a definitive article of “the eye-tracking data”. However, the claim initially introduces the “the eye-tracking information” and “the eye-tracking score”. There is insufficient antecedent basis for this limitation in the claim.
Claim 13 and 22 recites the limitation "the media item" and depends on claim 10, where such related definitive article was initially introduced as “digital media item”, “training media item”, “input media item”, and the “displayed training media item”. The recited limitation of “the media item” appears to be inconsistent and does not properly formulate which specific “media item” a segment corresponds to. There is insufficient antecedent basis for this limitation in the claim.
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claim(s) 1-23 is/are rejected under 35 U.S.C. 103 as being unpatentable over Bylinskii et al. (U.S. Doc. No. 11189066) in view of Jetley et al. (U.S. Pub. No. 20170308770).
Regarding claim 1, Bylinskii discloses a method for predicting an attention level of a viewer corresponding to one or more portions of a media item, comprising (col 3-4, “For instance, an importance map may be a graphical representation showing each pixel of the graphic object within a range from least important to most important where a higher importance indicates a higher probability that the pixel is viewed by a viewer.”; also, col 6, “the dynamic importance map 118a, through the color codes, may show predictions generated by the feedback engine 108 as to where viewers are expected to look at the input graphic object 116a.”): receiving the media item (col 11, “The method may begin at step 802, where the computer may receive an input graphic object. The input graphic object may be a graphic design or a data visualization. In some instances, the computer may display GUI with an upload tool for a user to upload an image, e.g., in a bitmap format, containing the input graphic object”); segmenting the media item into one or more segments, each segment corresponding to a portion of the media item (col 5, “The notion of importance may also depend upon higher level factors such as semantic categories of design elements (e.g., title text, axis text, data points) within a graphic object.”; also, col 7, “the smart thumbnail image 118c for a graphic object that is a data visualization may include title and other main supporting text, as well as data extremes—in case of a data table, top and bottom of the table, and in case of a data plot, left and right sides of the data plot.”); identifying one or more characteristics associated with each segment of the one or more segments (col 9, “An instruction to edit/update the input graphic object may include, for example, an instruction to resize a text or an image in the graphic object, an instruction to change font of a text in the graphic object, an instruction to change a color of a text or an image in the graphic object, and an instruction to move a text or an image in the graphic object.”; also, col 12, “the user may change font, size, or color of the text or rearrange the text or images. The user may also resize, crop, and/or change one or more colors of the picture.”); providing the segmented media item and identified characteristics to a machine-learning model (col 11, “At step 804, the computer may execute or deploy a neural network to generate an importance map for the input graphic object. The neural network may have been trained to optimize a loss function containing a pixel-level comparison of the importance output generated by the neural network and the ground truth dataset. In other words, the neural network may generate a pixel-level importance map for the input graphic object.”; also, col 7, “In some embodiments, the neural network 106 may be a convolutional neural network.”) heat map that predicts the attention level of the viewer corresponding to the one or more segments of the digital media item based on the segmented media item, identified characteristics, and the empirical data (col 11-12, “the importance map may be analogous to a heat map with different colors or shades of colors to represent the relative importance of various pixels within the input graphic object.”; also, col 11, “At step 806, the computer may display the importance map. The computer may display the importance map in the GUI side-by-side with the input graphic object. The importance map may be a graphical representation showing each pixel of the graphic object within a range from least important to most important”). Bylinskii does not disclosed that has been trained using empirical data from identified user gaze locations on training media items.
However, in a similar field of endeavor, Jetley discloses that has been trained using empirical data from identified user gaze locations on training media items (paragraph 57, “As illustrated in FIG. 3, in the exemplary embodiment, ground-truth saliency maps, referred to herein as attention maps 56, are constructed from the aggregated fixations of multiple observers, ignoring any temporal fixation information. Areas with a high fixation density are interpreted as receiving more attention.”; also, paragraph 58, “For example, at S202, eye gaze data is obtained for a set of observers for each training image (such as at least two observers). Eye gaze data may be obtained by observing the focus of the person's eye on an image, or simulated by receiving mouse clicks on the image made by observers. The collected eye gaze data is combined to generate a set of fixation coordinates 71.”; also, paragraph 40, “At S104, the neural network 64 is trained on the set of training images 62 and their associated attention maps 56 to predict saliency maps 12 for new images 14. The training includes learning weights for layers of the neural network through backward passes of the neural network”)
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified Bylinskii's invention of segmenting a graphic object into design elements, identifying per-segment characteristics, and applying a convolutional neural network to generate a heat-map importance prediction showing where viewers are expected to look on the graphic object with the features of Jetley's invention of training the neural network using empirical eye-fixation data aggregated from multiple human observers viewing the training images. One of ordinary skill in the art would have been motivated to make this combination because Jetley expressly teaches that "ground-truth saliency maps, referred to herein as attention maps 56, are constructed from the aggregated fixations of multiple observers," and Bylinskii itself acknowledges that bubble-view clicks and explicit importance-map crowdsourcing are "less expensive than tracking human eye gaze movements" but are proxies for the more accurate eye-gaze ground truth. First, Bylinskii expressly identifies eye-gaze measurements as the gold-standard ground truth that its proxies attempt to approximate, so substituting Jetley's aggregated multi-observer eye-fixation maps directly into Bylinskii's training pipeline yields the more accurate ground truth that Bylinskii itself recognizes as desirable. Second, Jetley demonstrates that as the number of observers increases, the AUC score of the aggregated fixation map approaches the inter-observer ceiling, providing an empirical basis for a person of ordinary skill to prefer eye-fixation ground truth over crowdsourced click proxies when sufficient eye-tracking data is available. Third, the substitution is a routine combination of known elements, namely, Bylinskii's convolutional-neural-network-based saliency prediction architecture for graphic designs and data visualizations, combined with Jetley's well-known multi-observer eye-fixation aggregation procedure, that yields the predictable result of a machine-learning model trained on empirical user-gaze locations for predicting attention on segmented media items.
Regarding claim 2, Bylinskii as modified by Jetley discloses the method of claim 1, further comprising displaying the heat map (Bylinskii: col 11, “At step 806, the computer may display the importance map. The computer may display the importance map in the GUI side-by-side with the input graphic object.”; col 11-12, “the importance map may be analogous to a heat map with different colors or shades of colors to represent the relative importance of various pixels within the input graphic object.”).
Regarding claim 3, Bylinskii as modified by Jetley the method of claim 1, wherein the heat map indicates a gaze of the viewer will be focused on each of the one or more segments of the media item (Bylinskii: col 6, “the dynamic importance map 118a, through the color codes, may show predictions generated by the feedback engine 108 as to where viewers are expected to look at the input graphic object 116a.”; also, col 3-4, “an importance map may be a graphical representation showing each pixel of the graphic object within a range from least important to most important where a higher importance indicates a higher probability that the pixel is viewed by a viewer.”).
Regarding claim 4, Bylinskii as modified by Jetley discloses the method of claims 1, wherein the one or more characteristics include at least one selected from size, location, text size, size rate between text and segment, center coordinates of segment, color, capital letters, non-alpha characters, number of characters, and a token term-frequency inverse document frequency (Bylinskii: col 9, “An instruction to edit/update the input graphic object may include, for example, an instruction to resize a text or an image in the graphic object, an instruction to change font of a text in the graphic object, an instruction to change a color of a text or an image in the graphic object, and an instruction to move a text or an image in the graphic object.”; also, col 12, “the user may change font, size, or color of the text or rearrange the text or images. The user may also resize, crop, and/or change one or more colors of the picture.”).
Regarding claim 5, Bylinskii as modified by Jetley discloses the method of claim 1, wherein the media item corresponds to at least one selected from a presentation, a website, a poster, and a document (Bylinskii: col 5, “a graphic object may be a graphic design, such as a poster or a data visualization (e.g., chart or a graph).”; also, col 7-8, “the first graphic object 202 may be a data visualization (e.g., containing a bar-graph), the second graphic object 204 may be a sales graphic (e.g., containing a sales graphic in a geographic region), and the third graphic object 206 may be graphic designs (e.g., containing a poster).”; also, col 12, “Within a webpage, for instance, the graphic object may have to be displayed within the main content, in the margins, or as banner advertisements at the top.”).
Regarding claim 6, Bylinskii as modified by Jetley discloses the method of claim 1, wherein each segment comprises at least one selected from a text item and an image item (Bylinskii: col 5, “The notion of importance may also depend upon higher level factors such as semantic categories of design elements (e.g., title text, axis text, data points) within a graphic object.”; also, col 12, “the user may change font, size, or color of the text or rearrange the text or images. The user may also resize, crop, and/or change one or more colors of the picture.”).
Regarding claim 7, Bylinskii as modified by Jetley discloses the method of claim 1, wherein the media item is received via a user interface, wherein the user interface includes a user interface control for uploading the media item (Bylinskii: col 11, “The method may begin at step 802, where the computer may receive an input graphic object. The input graphic object may be a graphic design or a data visualization. In some instances, the computer may display GUI with an upload tool for a user to upload an image, e.g., in a bitmap format, containing the input graphic object”).
Regarding claim 8, Bylinskii as modified by Jetley discloses the method of claim 1, wherein the user interface further comprises a first field configured to receive an indication of a number of pages included in the media item (Bylinskii: col 11, “The method may begin at step 802, where the computer may receive an input graphic object. The input graphic object may be a graphic design or a data visualization. In some instances, the computer may display GUI with an upload tool for a user to upload an image, e.g., in a bitmap format, containing the input graphic object”; also, col 6, “Retargeting is a common task for modern design tools because the input graphic object 116b may have to be targeted for different screen sizes, ranging from televisions to mobile devices with a plurality of screen dimensions. Furthermore, the input graphic object 116b may have to retargeted for other space/aspect ratio constraints, e.g., for different websites and different parts of webpages such as the main content portion or the margins.”).
Regarding claim 9, Bylinskii as modified by Jetley discloses the method of claim 1, wherein training the machine-learning model comprises (Bylinskii: col 11, “At step 804, the computer may execute or deploy a neural network to generate an importance map for the input graphic object. The neural network may have been trained to optimize a loss function containing a pixel-level comparison of the importance output generated by the neural network and the ground truth dataset.”): training media item includes one or more segments corresponding to a portion of the displayed training media item and each segment comprising one or more identifying characteristics (Bylinskii: col 5, “The notion of importance may also depend upon higher level factors such as semantic categories of design elements (e.g., title text, axis text, data points) within a graphic object.”, also, col 7, “the smart thumbnail image 118c for a graphic object that is a data visualization may include title and other main supporting text, as well as data extremes—in case of a data table, top and bottom of the table, and in case of a data plot, left and right sides of the data plot.”); training, based on the identified one or more identifying characteristics associated with each segment of the one or more segments (Bylinskii: col 14, “A fully convolutional neural network was trained using the ground truth importances Q.sub.iϵ[0,1] generated by one or more of the bubble-view and the explicit importance map methodology to optimize a sigmoid cross entropy loss for parameters Θ over all pixels i=1, 2, . . . , N. The loss function may be defined as follows:”) machine-learning model configured to receive an input media item and predict the attention level of a viewer corresponding to the one or more segments of the input media item (Bylinskii: col 11, “The method may begin at step 802, where the computer may receive an input graphic object. The input graphic object may be a graphic design or a data visualization”; also, col 6, “the dynamic importance map 118a, through the color codes, may show predictions generated by the feedback engine 108 as to where viewers are expected to look at the input graphic object 116a.”). Bylinskii does not disclose displaying, to a subject, a training media item on a display, receiving, from one or more sensors, eye-tracking information corresponding to a gaze location of the subject on the displayed training media item; determining an eye-tracking score for each of the one or more segments of the training media item based on the eye-tracking data; and the determined eye-tracking score for each of the one or more segments.
However, in a similar field of endeavor, Jetley discloses displaying, to a subject, a training media item on a display (paragraph 99, “The eye tracking data is collected using a head-mounted eye tracking device for 15 different viewers. The 1003 images of this dataset cover natural indoor and outdoor scenes.”; also, paragraph 56, “In an analysis of 300 images viewed by 39 observers, it was found that the fixations for a set of n observers match those from a different set of n observers with an AUC score that increases with the increase in the value of n.”, receiving, from one or more sensors, eye-tracking information corresponding to a gaze location of the subject on the displayed training media item (paragraph 58, “For example, at S202, eye gaze data is obtained for a set of observers for each training image (such as at least two observers). Eye gaze data may be obtained by observing the focus of the person's eye on an image, or simulated by receiving mouse clicks on the image made by observers. The collected eye gaze data is combined to generate a set of fixation coordinates 71.”; also, paragraph 99, “The eye tracking data is collected using a head-mounted eye tracking device for 15 different viewers”); determining an eye-tracking score for each of the one or more segments of the training media item based on the eye-tracking data (paragraph 59-60, “At S204, a binary fixation map 72, denoted b, is generated from the ground-truth eye-fixations 71. For each pixel i of the fixation map: At S206, the binary map b is then convolved (blurred) to approximate the general region fixated on. This can be done with a Gaussian kernel to produce a smoothed map 74, denoted y”); and the determined eye-tracking score for each of the one or more segments (paragraph 40, “At S104, the neural network 64 is trained on the set of training images 62 and their associated attention maps 56 to predict saliency maps 12 for new images 14. The training includes learning weights for layers of the neural network through backward passes of the neural network” also, paragraph 72, “An end-to-end learning method is adopted in which the fully-convolutional network 64 is trained on pairs of training images 62 and ground-truth saliency maps modeled as distributions p and g. The network 64 outputs predicted probability distributions p, while the ground-truth maps g are computed as described above for S102.”).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified Bylinskii's invention of segmenting a graphic object and applying a convolutional neural network to generate an importance heat map predicting where viewers are expected to look with the features of Jetley's invention of training the neural network by displaying training images to observers, receiving eye-tracking sensor data of the observers' gaze locations, determining per-pixel fixation scores aggregated across observers, and end-to-end learning the network on the resulting attention maps. One of ordinary skill in the art would have been motivated to make this combination because Jetley expressly teaches that "ground-truth saliency maps, referred to herein as attention maps 56, are constructed from the aggregated fixations of multiple observers," and that empirical eye-fixation data when aggregated across many observers approaches the inter-observer ceiling for saliency prediction accuracy. First, Bylinskii's bubble-view and explicit importance-map methodologies are expressly framed as efficient proxies for eye-gaze measurement, so substituting Jetley's direct eye-tracking sensor pipeline yields the more accurate ground truth that Bylinskii itself recognizes as the gold standard. Second, Jetley's per-pixel fixation map provides an immediately usable score for any segment defined by Bylinskii's design-element semantic categories simply by aggregating per-segment fixation density. Third, the combination uses only Bylinskii's convolutional neural network architecture and Jetley's well-known eye-fixation aggregation procedure to yield the predictable result of an attention-prediction model trained from empirical eye-tracking observations of segmented training media items.
Regarding claim 10, Bylinskii discloses a method for training a system for predicting attention level of a viewer corresponding to one or more portions of a digital media item, comprising (col 11, “At step 804, the computer may execute or deploy a neural network to generate an importance map for the input graphic object. The neural network may have been trained to optimize a loss function containing a pixel-level comparison of the importance output generated by the neural network and the ground truth dataset”; also, col 3-4, “an importance map may be a graphical representation showing each pixel of the graphic object within a range from least important to most important where a higher importance indicates a higher probability that the pixel is viewed by a viewer.”): displaying, to a subject, a training media item on a display, wherein the training media item includes one or more segments corresponding to a portion of the displayed training media item and each segment comprising one or more identifying characteristics (col 14, “In the bubble-view methodology, crowd workers may be shown a blurry image of a data visualization, and may be instructed to type a text caption describing the image. By clicking on different parts of the image, the crowd workers may reveal small regions—or bubbles—of the image at full resolution.”; also, col 5, “The notion of importance may also depend upon higher level factors such as semantic categories of design elements (e.g., title text, axis text, data points) within a graphic object.”, col 7, “the smart thumbnail image 118c for a graphic object that is a data visualization may include title and other main supporting text, as well as data extremes—in case of a data table, top and bottom of the table, and in case of a data plot, left and right sides of the data plot."); training, based on the identified one or more characteristics associated with each segment of the one or more segments (col 14, “A fully convolutional neural network was trained using the ground truth importances Q.sub.iϵ[0,1] generated by one or more of the bubble-view and the explicit importance map methodology to optimize a sigmoid cross entropy loss for parameters Θ over all pixels i=1, 2, . . . , N. The loss function may be defined as follows:”; col 7, “In some embodiments, the neural network 106 may be a convolutional neural network.”) machine-learning model configured to receive an input media item and predict the attention level of a viewer corresponding to the one or more segments of the input media item (col 11, “The method may begin at step 802, where the computer may receive an input graphic object. The input graphic object may be a graphic design or a data visualization.”; also, col 6, “the dynamic importance map 118a, through the color codes, may show predictions generated by the feedback engine 108 as to where viewers are expected to look at the input graphic object 116a”). Bylinskii does not disclose receiving, from one or more sensors, eye-tracking information corresponding to a gaze location of the subject on the displayed training media item; determining an eye-tracking score for each of the one or more segments of the training media item based on the eye-tracking data, and the determined eye-tracking score for each of the one or more segments.
However, in a similar field of endeavor, Jetley discloses receiving, from one or more sensors, eye-tracking information corresponding to a gaze location of the subject on the displayed training media item (paragraph 58, “For example, at S202, eye gaze data is obtained for a set of observers for each training image (such as at least two observers). Eye gaze data may be obtained by observing the focus of the person's eye on an image, or simulated by receiving mouse clicks on the image made by observers. The collected eye gaze data is combined to generate a set of fixation coordinates 71.”; also, paragraph 99, "The eye tracking data is collected using a head-mounted eye tracking device for 15 different viewers.”); determining an eye-tracking score for each of the one or more segments of the training media item based on the eye-tracking data (paragraph 59-60, “At S204, a binary fixation map 72, denoted b, is generated from the ground-truth eye-fixations 71. For each pixel i of the fixation map: "At S206, the binary map b is then convolved (blurred) to approximate the general region fixated on. This can be done with a Gaussian kernel to produce a smoothed map 74, denoted y.”), and the determined eye-tracking score for each of the one or more segments (paragraph 63-64, “At S210, the normalized attention map 76 is converted to an attention map probability distribution 56, denoted g, with a softmax function: where g.sub.i is the probability of pixel i being fixated on”; also, paragraph 72, “An end-to-end learning method is adopted in which the fully-convolutional network 64 is trained on pairs of training images 62 and ground-truth saliency maps modeled as distributions p and g.”).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified Bylinskii's invention of training a convolutional neural network on segmented graphic objects with per-segment characteristics to predict an importance heat map of where viewers are expected to look with the features of Jetley's invention of collecting eye-tracking sensor data from multiple observers viewing the training images and aggregating that gaze data into a per-pixel eye-tracking score that becomes the ground-truth training signal for the neural network. One of ordinary skill in the art would have been motivated to make this combination because Jetley expressly teaches that "ground-truth saliency maps, referred to herein as attention maps 56, are constructed from the aggregated fixations of multiple observers," and Bylinskii itself recognizes that bubble-view and explicit-importance-map crowdsourcing methodologies are "less expensive than tracking human eye gaze movements" but operate as proxies for the more accurate eye-gaze ground truth. First, Bylinskii expressly identifies human eye gaze measurements as the gold-standard ground truth that its proxies attempt to approximate, providing explicit motivation to substitute Jetley's aggregated multi-observer eye-fixation scores into Bylinskii's training pipeline. Second, Jetley's per-pixel fixation map provides exactly the pixel-level importance ground truth that Bylinskii's training loss function consumes, allowing direct substitution without any architectural change. Third, the substitution is a routine combination of known elements, namely, Bylinskii's segmentation and per-segment characteristic-aware convolutional neural network training architecture combined with Jetley's well-known multi-observer eye-tracking sensor data aggregation procedure, that yields the predictable result of a machine-learning model trained on empirical user-gaze data of segmented training media items.
Regarding claim 11, Bylinskii as modified by Jetley discloses the method of claim 10, repeating the displaying, receiving, determining, and training with a second training media item.
However, in a similar field of endeavor, Jetley further discloses repeating the displaying, receiving, determining, and training with a second training media item (paragraph 58, “For example, at S202, eye gaze data is obtained for a set of observers for each training image (such as at least two observers). Eye gaze data may be obtained by observing the focus of the person's eye on an image, or simulated by receiving mouse clicks on the image made by observers. The collected eye gaze data is combined to generate a set of fixation coordinates 71.”; also, paragraph 72, “An end-to-end learning method is adopted in which the fully-convolutional network 64 is trained on pairs of training images 62 and ground-truth saliency maps modeled as distributions p and g.”).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified Bylinskii's invention of training a convolutional neural network on segmented graphic objects with per-segment characteristics to predict an importance heat map with the features of Jetley's invention of iterating the per-image gaze-data collection and end-to-end network training across each training image of a set of training images. One of ordinary skill in the art would have been motivated to make this combination because Jetley expressly teaches that "eye gaze data is obtained for a set of observers for each training image," and that "the fully-convolutional network 64 is trained on pairs of training images 62 and ground-truth saliency maps modeled as distributions p and g," yielding a training pipeline that operates on a set of training images one image at a time. First, iterating the display, gaze-data collection, fixation-score determination, and network training cycle across multiple training images is recognized as the standard practice for training convolutional neural networks, because per-image fixation maps are insufficient to constrain the network's parameters on their own. Second, Jetley's recitation of "for each training image" and "pairs of training images" makes clear that the cycle is intended to operate over a set of training images rather than a single image. Third, the combination yields the predictable result of a training pipeline in which the displaying, receiving, determining, and training operations are repeated for a second training media item.
Regarding claim 12, Bylinskii as modified by Jetley discloses the method of claim 10, further comprising repeating the displaying, receiving, determining, and training with a second subject (Bylinskii: col 14, “In the bubble-view methodology, crowd workers may be shown a blurry image of a data visualization, and may be instructed to type a text caption describing the image. By clicking on different parts of the image, the crowd workers may reveal small regions—or bubbles—of the image at full resolution.”; also, col 14, “In the explicit importance maps methodology, crowd workers may be provided with images with graphic designs and asked to label important regions of designs using binary masks. By averaging the responses, ground truth importance maps of the graphic design may be generated to train and/or test a neural network.”).
Regarding claim 13, Bylinskii as modified by Jetley discloses the method of claim 10, further comprising segmenting the training media item into one or more segments, each segment corresponding to a portion of the media item (Bylinskii: col 5, “The notion of importance may also depend upon higher level factors such as semantic categories of design elements (e.g., title text, axis text, data points) within a graphic object”; also, col 7, “the smart thumbnail image 118c for a graphic object that is a data visualization may include title and other main supporting text, as well as data extremes—in case of a data table, top and bottom of the table, and in case of a data plot, left and right sides of the data plot.”).
Regarding claim 14, Bylinskii as modified by Jetley discloses the method of claim 10, further comprising identifying one or more characteristics associated with each segment of the one or more segments (Bylinskii: col 9, “An instruction to edit/update the input graphic object may include, for example, an instruction to resize a text or an image in the graphic object, an instruction to change font of a text in the graphic object, an instruction to change a color of a text or an image in the graphic object, and an instruction to move a text or an image in the graphic object.”; also, col 12, “the user may change font, size, or color of the text or rearrange the text or images. The user may also resize, crop, and/or change one or more colors of the picture.”).
Regarding claim 15, Bylinskii as modified by Jetley discloses the method of claim 10, eye-tracking score is based on an amount of time the gaze location of the subject is focused on the respective segment.
However, in a similar field of endeavor, Jetley discloses wherein the eye-tracking score is based on an amount of time the gaze location of the subject is focused on the respective segment (paragraph 59-60, “At S204, a binary fixation map 72, denoted b, is generated from the ground-truth eye-fixations 71. For each pixel i of the fixation map: At S206, the binary map b is then convolved (blurred) to approximate the general region fixated on. This can be done with a Gaussian kernel to produce a smoothed map 74, denoted y.”; also, paragraph 64, “where g.sub.i is the probability of pixel i being fixated on”).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified Bylinskii's invention of training a convolutional neural network to predict an importance heat map for a segmented graphic object with the features of Jetley's invention of computing a per-pixel eye-tracking score from the aggregated count of discrete fixation events captured at each pixel of the training image. One of ordinary skill in the art would have been motivated to make this combination because Jetley expressly teaches that "a binary fixation map 72, denoted b, is generated from the ground-truth eye-fixations 71," with the entries of the map recording per-pixel fixation events that are then convolved to capture per-region fixation density. First, each fixation event captured by Jetley's binary fixation map constitutes a minimum dwell interval of the observer's gaze on the corresponding pixel, so the aggregated count of fixation events at a given pixel or region is directly proportional to the amount of time the subject's gaze was focused on that pixel or region. Second, Jetley's Gaussian convolution step explicitly converts the per-pixel fixation events into a smoothed per-region score, providing a per-segment eye-tracking score whose magnitude tracks the aggregate dwell time at the segment. Third, the combination yields the predictable result of an eye-tracking score whose value at each segment is based on the amount of time the subject's gaze is focused on that segment.
Regarding claim 16, Bylinskii as modified by Jetley discloses the method of claim 10, wherein the system is configured to output a heat map that predicts a location of a gaze of the viewer's gaze on the displayed media item based on the eye-tracking score for each of the one or more segments and the one or more identifying characteristics for each of the one or more segments (Bylinskii: col 11-12, “the importance map may be analogous to a heat map with different colors or shades of colors to represent the relative importance of various pixels within the input graphic object.”; also, col 6, “the dynamic importance map 118a, through the color codes, may show predictions generated by the feedback engine 108 as to where viewers are expected to look at the input graphic object 116a.”; also, col 11, “At step 806, the computer may display the importance map. The computer may display the importance map in the GUI side-by-side with the input graphic object.”).
Regarding claim 17, Bylinskii as modified by Jetley discloses the method of claim 16, heat map indicates a length of time that a gaze of the viewer's gaze will be focused on each segment of the one or more segments.
However, in a similar field of endeavor, Jetley discloses wherein the heat map indicates a length of time that a gaze of the viewer's gaze will be focused on each segment of the one or more segments (paragraph 59-60, “At S204, a binary fixation map 72, denoted b, is generated from the ground-truth eye-fixations 71. For each pixel i of the fixation map: At S206, the binary map b is then convolved (blurred) to approximate the general region fixated on. This can be done with a Gaussian kernel to produce a smoothed map 74, denoted y.”; also, paragraph 64, “where g.sub.i is the probability of pixel i being fixated on”).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified Bylinskii's invention of generating an importance heat map for a segmented graphic object with the features of Jetley's invention of representing each pixel of the heat map by an aggregated count of fixation events at that pixel, where each fixation event constitutes a minimum dwell interval of the viewer's gaze. One of ordinary skill in the art would have been motivated to make this combination because Jetley expressly teaches that "a binary fixation map 72, denoted b, is generated from the ground-truth eye-fixations 71," and that the resulting per-pixel value "gi is the probability of pixel i being fixated on," such that the heat-map value at each pixel directly reflects how often the viewer's gaze landed at that pixel. First, each discrete fixation event captured by Jetley's per-pixel scoring is a minimum dwell interval, so the heat-map value at a given segment is directly proportional to the aggregate length of time the viewer's gaze is focused on that segment. Second, encoding aggregate gaze duration into a per-pixel heat map yields a more informative attention prediction than a binary attended-or-not output. Third, the combination yields the predictable result of a heat map whose value at each segment indicates a length of time the viewer's gaze will be focused on that segment.
Regarding claim 18, Bylinskii as modified by Jetley discloses the method of claim 10, wherein the eye-tracking data includes an eye-tracking bubble, wherein the eye-tracking bubble indicates a portion of the media item corresponding to the subject's gaze location (Bylinskii: col 14, “In the bubble-view methodology, crowd workers may be shown a blurry image of a data visualization, and may be instructed to type a text caption describing the image. By clicking on different parts of the image, the crowd workers may reveal small regions—or bubbles—of the image at full resolution. These clicks may be referred to as bubble clicks. The regions with most amount of bubble clicks may be the more important regions within the image of data visualization.”; also, Abstract, “The one or more neural networks are trained to optimize a loss function comprising a pixel-level comparison between the outputs generated by the neural networks and the ground truth dataset generated from a bubble view methodology or an explicit importance maps methodology.”).
Regarding claim 19, Bylinskii as modified by Jetley discloses the method of claim 10, wherein each segment comprises at least one selected from a word and an image (Bylinskii: col 5, “The notion of importance may also depend upon higher level factors such as semantic categories of design elements (e.g., title text, axis text, data points) within a graphic object.”; also, col 12, “the user may change font, size, or color of the text or rearrange the text or images. The user may also resize, crop, and/or change one or more colors of the picture.”).
Regarding claim 20, Bylinskii as modified by Jetley discloses the method of claim 10, wherein the one or more characteristics include at least one selected from size, location, text size, size rate between text and segment, center coordinates of segment, color, capital letters, non-alpha characters, number of characters, and a token term-frequency inverse document frequency (Bylinskii: col 9, “An instruction to edit/update the input graphic object may include, for example, an instruction to resize a text or an image in the graphic object, an instruction to change font of a text in the graphic object, an instruction to change a color of a text or an image in the graphic object, and an instruction to move a text or an image in the graphic object.”; also, col 12, “the user may change font, size, or color of the text or rearrange the text or images.”).
Regarding claim 21, Bylinskii as modified by Jetley discloses the method of claim 10, wherein the media item corresponds to at least one selected from a presentation, a website, a poster, and a document (Bylinskii: col 5, “a graphic object may be a graphic design, such as a poster or a data visualization (e.g., chart or a graph).”; also, col 7-8, “the first graphic object 202 may be a data visualization (e.g., containing a bar-graph), the second graphic object 204 may be a sales graphic (e.g., containing a sales graphic in a geographic region), and the third graphic object 206 may be graphic designs (e.g., containing a poster).”; also, col 12, “Within a webpage, for instance, the graphic object may have to be displayed within the main content, in the margins, or as banner advertisements at the top”).
Regarding claim 22, Bylinskii as modified by Jetley discloses the method of claim 10, further comprising identifying one or more characteristics of the media item by applying natural language processing to the segmented media item (Bylinskii: col 5, “The notion of importance may also depend upon higher level factors such as semantic categories of design elements (e.g., title text, axis text, data points) within a graphic object.”, also, col 7, “the smart thumbnail image 118c for a graphic object that is a data visualization may include title and other main supporting text, as well as data extremes—in case of a data table, top and bottom of the table, and in case of a data plot, left and right sides of the data plot”).
Regarding claim 23, Bylinskii discloses a system for predicting an attention level of a viewer corresponding to one or more portions of a media item, comprising (col 3-4, “an importance map may be a graphical representation showing each pixel of the graphic object within a range from least important to most important where a higher importance indicates a higher probability that the pixel is viewed by a viewer.”): a display; one or more processors; a memory; and one or more programs, wherein the one or more programs are stored in the memory and configured to be executed by the one or more processors, the one or more programs including instructions for (col 6, “The processor 102 may execute multiple engines (e.g., a real-time feedback engine 108, a retargeting engine 110, a smart thumbnails engine 112, a color palette (or theme) extraction engine 114) that may, in turn, execute a neural network 106 stored in the non-transitory memory 104.”; also, col 11, “At step 806, the computer may display the importance map. The computer may display the importance map in the GUI side-by-side with the input graphic object”): receiving the media item (col 11, “The method may begin at step 802, where the computer may receive an input graphic object. The input graphic object may be a graphic design or a data visualization. In some instances, the computer may display GUI with an upload tool for a user to upload an image, e.g., in a bitmap format, containing the input graphic object.”); segmenting the media item into one or more segments, each segment corresponding to a portion of the media item (col 5, “The notion of importance may also depend upon higher level factors such as semantic categories of design elements (e.g., title text, axis text, data points) within a graphic object”; also, col 7, “the smart thumbnail image 118c for a graphic object that is a data visualization may include title and other main supporting text, as well as data extremes—in case of a data table, top and bottom of the table, and in case of a data plot, left and right sides of the data plot.”); identifying one or more characteristics associated with each segment of the one or more segments (col 9, “An instruction to edit/update the input graphic object may include, for example, an instruction to resize a text or an image in the graphic object, an instruction to change font of a text in the graphic object, an instruction to change a color of a text or an image in the graphic object, and an instruction to move a text or an image in the graphic object.”); providing the segmented media item and identified characteristics to a machine- learning model (col 11, “At step 804, the computer may execute or deploy a neural network to generate an importance map for the input graphic object. The neural network may have been trained to optimize a loss function containing a pixel-level comparison of the importance output generated by the neural network and the ground truth dataset.”; also, col 7, “In some embodiments, the neural network 106 may be a convolutional neural network.”) a heat map that predicts the attention level of the viewer corresponding to the one or more segments of the digital media item based on the segmented media item, identified characteristics, and the empirical data (col 11-12, “the importance map may be analogous to a heat map with different colors or shades of colors to represent the relative importance of various pixels within the input graphic object.”; also, col 6, “the dynamic importance map 118a, through the color codes, may show predictions generated by the feedback engine 108 as to where viewers are expected to look at the input graphic object 116a.”). Bylinskii does not disclose that has been trained using empirical data from identified user gaze locations on training media items.
However, in a similar field of endeavor, Jetley discloses that has been trained using empirical data from identified user gaze locations on training media items (paragraph 57, “As illustrated in FIG. 3, in the exemplary embodiment, ground-truth saliency maps, referred to herein as attention maps 56, are constructed from the aggregated fixations of multiple observers, ignoring any temporal fixation information."”; also, paragraph 40, “At S104, the neural network 64 is trained on the set of training images 62 and their associated attention maps 56 to predict saliency maps 12 for new images 14.”).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified Bylinskii's invention of a system having a processor, memory, and instructions for segmenting a graphic object, identifying per-segment characteristics, and using a convolutional neural network to generate an importance heat map showing viewer-attention predictions with the features of Jetley's invention of training the neural network on aggregated multi-observer eye-fixation maps of training images. One of ordinary skill in the art would have been motivated to make this combination because Jetley expressly teaches that the "network 64 is trained on the set of training images 62 and their associated attention maps 56 to predict saliency maps 12 for new images 14," and Bylinskii itself recognizes that human eye-gaze data is the gold-standard ground truth that its crowdsourced alternatives attempt to approximate. First, Bylinskii's hardware platform, namely, a processor, non-transitory memory, and neural network instructions, is directly compatible with Jetley's training pipeline because both operate end-to-end on the same convolutional neural network architecture. Second, Jetley's per-observer fixation maps, when aggregated and convolved with a Gaussian kernel, yield exactly the pixel-level importance ground truth that Bylinskii's training loss function consumes, allowing direct substitution without architectural change. Third, the combination yields the predictable result of a system that segments an input media item and generates an attention-prediction heat map, where the underlying machine-learning model was trained from empirical user-gaze data.
Conclusion
Any inquiry concerning this communication or earlier communications from the examiner should be directed to Jai Li whose telephone number is (571)272-1170. The examiner can normally be reached Mon-Thu between 06:00-16:00 EST.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Xiao Wu can be reached at (571)272-7761. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/JAI W LI/Junior Examiner, Art Unit 2613
/XIAO M WU/Supervisory Patent Examiner, Art Unit 2613