Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
DETAILED ACTION
Claim Rejections - 35 USC § 112
The following is a quotation of 35 U.S.C. 112(b):
(b) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention.
The following is a quotation of 35 U.S.C. 112 (pre-AIA ), second paragraph:
The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the applicant regards as his invention.
Claims 1-20 rejected under 35 U.S.C. 112(b) or 35 U.S.C. 112 (pre-AIA ), second paragraph, as being indefinite for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor, or for pre-AIA the applicant regards as the invention. Claims 1, 11, 20 recite “other audio features and other audio descriptions” the recitation of “other audio features,” is considered to lack clear antecedent and render the claims indefinite; claims 6, 16 recite “other audio features,” an a similarly indefinite manner. Claims 2-10, 12-19 do not remedy and are similarly rejected. Claims 2, 12 additionally recite “visual features,” in a manner lacking clear antecedent, presumably the recited “a plurality of extracted video features” of the parent claims. Claim 5, 15 additionally recite “the plurality of predetermined audio detail levels,” in a manner lacking clear antecedent, presumably the recited “a plurality of predetermined audio description detail levels” of claims 4, 14. Claims 7, 17 appear to contain a typographical error and include the recitation “the video content ant the at least one characteristic of the user,” Examiner will presume the claims recite “and.” Claim 9, 19 recite generating a similarity measure based on “characteristics of the user,” Parent claims 1, 7, 11, 17 recite “at least one characteristic of a user,” which is met by a singular characteristic. Appropriate correction is required.
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102 of this title, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
The factual inquiries set forth in Graham v. John Deere Co., 383 U.S. 1, 148 USPQ 459 (1966), that are applied for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows:
1. Determining the scope and contents of the prior art.
2. Ascertaining the differences between the prior art and the claims at issue.
3. Resolving the level of ordinary skill in the pertinent art.
4. Considering objective evidence present in the application indicating obviousness or nonobviousness.
Claims 1-6, 10-16, 20 rejected under 35 U.S.C. 103 as being unpatentable over Mayhar: 10999566, hereinafter Ma further in view of Pavel: Rescribe: Authoring and Automatically Editing Audio Descriptions (copy provided by Examiner, copyright 2020, and hereinafter Pa).
Regarding claim 1
Ma teaches:
A method, in a data processing system, for generating audio descriptions of visual elements of video content (Ma: Abstract; Col 3:8-3:14: system generates descriptions of visual content, presents said descriptions audibly), the method comprising:
receiving a plurality of extracted video features for the video content from an image recognition and analysis computing system that performs image recognition operations to extract video features from the video content (Ma: Col 4:10-4:24, 10:4-10:23, 11:12-11:14; Fig 1, 3: a content processing engine, operable to perform image recognition, video processing, etc. on video frames detects features therein such as objects, faces, etc. consistent across a plurality of frames; actions transiting over a plurality of frames; etc.; said processing engine operative to output vector data; said vectors received by a description generation engine operable to form textual descriptions based thereon);
generating, based on the extracted video features, one or more audio descriptions that describe the extracted video features audibly (Ma: Col 4:10-4:24, 5:63-6:3; Fig 1, 3: system generates textual descriptions from feature vectors extracted from some or all frames);
determining at least one temporal location within the video content in which to place the one or more audio descriptions, wherein the at least one temporal location is determined based on a criterion to minimize overlap of the one or more audio descriptions with other audio features of the video content (Ma: Col 12:28-12:59, 14:65-15:10, 16:26-16:30; Fig 3, 4: system determines placement of descriptive audio to correspond with playback of the relevant video segment and based on presence of extant audio therein, sound levels thereof);
generating an audio description data structure based on the one or more audio descriptions and the at least one temporal location within the video content for the one or more audio descriptions (Ma: Col 8:50-8:59, 13:10-13:24; Fig 3, 4: system produces textual description files, associated with, correlated upon, mapped to, etc. particular segment start and segment end timestamps from which audio is rendered in appropriate positions); and
providing the audio description data structure to a client computing device for playback of the video content and output of the one or more audio descriptions during the playback of the video content in accordance with the audio description data structure (Ma: Col 12:28-12:59: text description transmitted in audio form to user for output during playback).
Ma teaches the minimization of overlap with other audio objects by detection thereof (Ma: Col 13:60-14:37) and thus strongly suggests but does not explicitly teach the minimization of overlap with respect to “other audio descriptions.”
In a related field of endeavor Pa teaches a system and method for authoring and editing audio descriptions of visual elements of video content (Pa: Abstract; Fig 1), the method comprising:
determining at least one temporal location within the video content in which to place the one or more audio descriptions, wherein the at least one temporal location is determined based on a criterion to minimize overlap of the one or more audio descriptions with other audio features of the video content and other audio descriptions (Pa: Identifying description locations; pp 5; Overlap score, pp 6: system identifies appropriate regions in source audio within which to place audio descriptions to thereby minimize overlap and reduce overall disruption such as by applying “penalty for overlapping with speech in the source track, or a previously placed audio description,”);
generating an audio description data structure based on the one or more audio descriptions and the at least one temporal location within the video content for the one or more audio descriptions (Pa: Scoring audio description compositions; pp 5: system represents the composition as a data structure comprising a series of tuples describing start time and duration for each of a plurality of descriptions).
It would have been obvious to one of ordinary skill in the art before the effective filing date of the instant application to improve the Ma system and method by utilization or inclusion of the Pa data structure to manage description insertion for at least the purpose of minimizing the likelihood of inserted description overlap; one of ordinary skill in the art would have expected only predictable results therefrom.
Regarding claim 2
Ma in view of Pa teaches or suggests:
The method of claim 1, wherein the one or more audio descriptions are audio descriptions comprising descriptive audio content for presentation to blind and visually impaired (BVI) persons to describe visual features of the video content that are not able to be perceived by the BVI persons (Ma: Col 3:8-3:10: system presents textural descriptions of on screen video to “aid visually impaired users,”); (Pa: Introduction, pp 1: “Audio descriptions (AD) make videos accessible by describing important visual content in the audio for those who cannot see it,”). The claim is considered obvious over Ma as modified by Pa as addressed in the base claim as it would have been obvious to apply the further teaching of Ma and/or Pa to the modified device of Ma and Pa; one of ordinary skill in the art would have expected only predictable results therefrom.
Regarding claim 3
Ma in view of Pa teaches or suggests:
The method of claim 1, wherein generating one or more audio descriptions comprises: retrieving a user profile corresponding to a user requesting generating of the one or more audio descriptions for the video content, wherein the user profile specifies an audio description detail level setting specifying a level of detail to be included in the one or more audio descriptions; and generating the one or more audio descriptions based on the audio description detail level setting (Ma: Col 11:47-11:63: system determines preferences of an active user including description settings thereof to determine custom settings for the textual descriptions output to a user). The claim is considered obvious over Ma as modified by Pa as addressed in the base claim as it would have been obvious to apply the further teaching of Ma and/or Pa to the modified device of Ma and Pa; one of ordinary skill in the art would have expected only predictable results therefrom.
Regarding claim 4
Ma in view of Pa teaches or suggests:
The method of claim 3, wherein the audio description detail level setting is one of a plurality of predetermined audio description detail levels, and wherein each predetermined audio description detail level comprises a different amount of detail from other predetermined audio description detail levels with regard to descriptions of video features to be included in audio descriptions (Ma: Col 3:52-3:64, 11:47-11:63, 15:12-15:20: system comprises plural settings for textual descriptions at different levels of granularity); (Pa: Generating candidate descriptions, pp 5; Scoring audio description compositions, pp 5; Coherence score, pp 5, 6: system generates ranked plurality of candidate description comprising different levels of detail). The claim is considered obvious over Ma as modified by Pa as addressed in the base claim as it would have been obvious to apply the further teaching of Ma and/or Pa to the modified device of Ma and Pa; one of ordinary skill in the art would have expected only predictable results therefrom.
Regarding claim 5
Ma in view of Pa teaches or suggests:
The method of claim 4, wherein the plurality of predetermined audio detail levels (Ma: Col 3:52-3:64, 4:22-4:35, 10:10-10:14, 11:47-11:63, 15:12-15:20: such as granularity levels of parameters, characteristics, etc.); (Pa: Generating candidate descriptions, pp 5; Scoring audio description compositions, pp 5; Coherence score, pp 5, 6: system generates ranked plurality of candidate description comprising different levels of detail) comprises:
a first predetermined audio description detail level comprises identifiers of the video features, but no location information specifying a relative location of the video features to one another, and no descriptor terms associated with the video features (Ma: id.: system comprises plural settings for textual descriptions at different levels of granularity, verbosity, etc.; such as reductions of detail) (Pa: id.: description comprising locational relations of the features, other descriptors are scored, and the relations/locations, above, below; and descriptors are selectably dropped, such as by dropping adjectives, particular phrases, etc.—“people walking with an sky,”), a second predetermined audio description detail level comprises the identifiers of the video features and descriptor terms associated with the video features, but no location information (Pa: id.: “people walking along a beach,” comprises descriptors but no relations thereamong), and a third predetermined audio description detail level comprises the identifiers of the video features, the location information, and descriptor terms associated with the video features (Pa: id.: the original description). In Ma the system selectably progresses along levels of refinement such as those claimed as appropriate to a particular level of granularity stripping relational, location type information and descriptors accordingly, such as in relation to the Pa taught removal of relational phrases, adjective descriptors, etc. It would have been obvious to one of ordinary skill in the art before the effective filing date of the instant application as a matter of design choice to assign granularity or verbosity values as taught or suggested by Ma to the dropping of descriptor terms and relational terms as taught or suggested by Pa such as to avoid over-describing terms inferable from the video, etc.; abbreviating or elongating descriptions to fit within available segments or in concert with stylistic choices; etc.; one of ordinary skill in the art would have expected only predictable results therefrom.
Regarding claim 6
Ma in view of Pa teaches or suggests:
The method of claim 4, wherein determining the at least one temporal location within the video content comprises iteratively generating audio descriptions at different predetermined audio description detail levels until an audio description having a temporal length that fits within a rendering window, with a predetermined level of acceptable overlap with other audio features and other audio descriptions, is generated (Ma: Abstract; Col14:39-14:52: system iteratively lengthens, shortens insertion based on available length, granularity parameters, etc.); (Pa: Abstract; Identifying description locations; pp 5; Overlap score, pp 6: system identifies appropriate regions in source audio within which to place audio descriptions to thereby minimize overlap and reduce overall disruption such as by applying “penalty for overlapping with speech in the source track, or a previously placed audio description,”. Please see claim 5 supra—it would have been obvious to one of ordinary skill in the art before the effective filing date of the instant application as a matter of design choice to assign granularity or verbosity values as taught or suggested by Ma to the dropping of descriptor terms and relational terms as taught or suggested by Pa such as to avoid over-describing terms inferable from the video, etc.; abbreviating or elongating descriptions to fit within available segments or in concert with stylistic choices; etc.; one of ordinary skill in the art would have expected only predictable results therefrom.
Regarding claim 10
Ma in view of Pa teaches or suggests:
The method of claim 1, wherein the plurality of extracted video features is a filtered set of extracted video features having fewer extracted video features than an original set of extracted video features extracted from the video content by the image recognition and analysis system , and wherein the filtered set of extracted video features is generated by filtering the original set of extracted video features in accordance with one or more user specified filter criteria that specify types of extracted features for which audio descriptions are not to be generated (Ma: Col 3:52-3:64, 11:47-11:63, 15:12-15:20: system comprises filters, reduces, etc. descriptions based on granularity settings); (Pa: Generating candidate descriptions, pp 5; Informativeness score, pp 6: system selectable excludes particular language, types of language, etc. from the description). The claim is considered obvious over Ma as modified by Pa as addressed in the base claim as it would have been obvious to apply the further teaching of Ma and/or Pa to the modified device of Ma and Pa; one of ordinary skill in the art would have expected only predictable results therefrom.
Regarding claim 11, 20—the claims are considered to recite substantially similar subject matter to that of claim 1 and are similarly rejected.
Regarding claim 12—the claim is considered to recite substantially similar subject matter to that of claim 2 and is similarly rejected.
Regarding claim 13—the claim is considered to recite substantially similar subject matter to that of claim 3 and is similarly rejected.
Regarding claim 14—the claim is considered to recite substantially similar subject matter to that of claim 4 and is similarly rejected.
Regarding claim 15—the claim is considered to recite substantially similar subject matter to that of claim 5 and is similarly rejected.
Regarding claim 16—the claim is considered to recite substantially similar subject matter to that of claim 6 and is similarly rejected.
Claims 7-9, 17-19 rejected under 35 U.S.C. 103 as being unpatentable over Mayhar: 10999566, hereinafter Ma further in view of Pavel: Rescribe: Authoring and Automatically Editing Audio Descriptions (copy provided by Examiner, copyright 2020, and hereinafter Pa) as applied to claims 1-6, 10-16, 20 supra and further in view of well-known search of data structures as evidenced by Miller: 11238899 hereinafter Mi.
Regarding claim 7
Ma in view of Pa teaches or suggests:
The method of claim 1, further comprising:
performing a search of audio description data structures based on an identification of the video content (Ma: Col 5:47-5:50, 6:24-6:33, etc.: system provides search to located content including particular scenes based on descriptions thereof) and at least one characteristic of a user requesting generation of the one or more audio descriptions (Ma: Col 11:45-12:4, 12:14-12:20: system searches in concert with user profile data to resolve particular descriptions) to thereby identify a matching audio description data structure corresponding to the video content and provide the matching audio description data structure to the client computing device as the audio description data structure such as for playback (Ma Col 12:40-12:63, 14:10-10:24: determined audio descriptions comprise a structure associated same with temporal locations, renderable artifacts, and operable in concert with audio, video content to generate new descriptions for output).
Ma in view of Pa does not explicitly discuss the system operable to search a stored audio description data structures in a storage system to thereby identify a matching audio description data structure corresponding to the video content and at least one characteristic of the user; find the matching audio description data structure in the storage system of previously stored description structures operable to query and retrieve a matching audio description data structure based on an identification of the video content and at least one characteristic of a user requesting generation of the one or more audio description and operable for provide the matching audio description data structure to the client computing device as the audio description data structure.
Examiner takes official notice that searching of a system such as that of Ma in view of Pa using the claimed search dynamics was well known in the art before the effective filing date of the instant application and would have comprised an obvious inclusion for at least the purpose of using conventional data structures to determine and utilize audio descriptions in concert with video media. As evidence consider Mi teaches a system and method for generating an audio description media file (Mi: Abstract) wherein the system comprises stored audio description data structures in a storage system comprising audio description data and temporal associated data thereof for synchrony of audio with a video time (Mi: Abstract; Col 7:39-7:67: system comprises media file storage with associated records comprising descriptions maintained therein); such as for performing a search of stored audio description data structures in a storage system based on an identification of the video content (Mi: Col 34:33-34:47; Fig 1: media stored with unique identifier for indexing and subsequent reference), and at least one characteristic of a user (Mi: Col 35:53-36:8: customer and user configuration settings operable in association with a media file, media project, etc.), to identify a matching audio description data structure corresponding to the video content ant <sic> the at least one characteristic of the user (Mi: Col 8:41-8:57, 35:19-35:35: system operable to perform search based on indexes of the stored information, such as for determining appropriate files); and in response to finding the matching audio description data structure in the storage system, retrieving the matching audio description data structure (Mi: Col 11:1-11:8; Fig 13: system stores the description data resulting from a match for subsequent retrieval) and providing the matching audio description data structure to the client computing device as the audio description data structure (Mi: Col 8:27-8:40: system provides determined information to user, client computer thereof). It would have been obvious to one of ordinary skill in the art before the effective filing date of the instant application to utilize the Mi disclosed data structures and search dynamics to traverse, explore and resolve data within the Ma in view of Pa system and method; one of ordinary skill in the art would have expected only predictable results therefrom.
Regarding claim 8
Ma in view of Pa in view of Mi teaches or suggests:
The method of claim 7, wherein the at least one characteristic of the user comprises one or more of an identifier of a visual impairment of the user or a specified level of detail for inclusion in audio descriptions of video features (Ma: Col 3:52-3:64, 11:47-11:63, 15:12-15:20: system comprises plural settings for textual descriptions at different levels of granularity); (Pa: Generating candidate descriptions, pp 5; Scoring audio description compositions, pp 5; Coherence score, pp 5, 6: system generates ranked plurality of candidate description comprising different levels of detail). The claim is considered obvious over Ma as modified by Pa and Mi as addressed in the base claim as it would have been obvious to apply the further teaching of Ma, Pa, and/or Mi to the modified device of Ma, Pa, and Mi; one of ordinary skill in the art would have expected only predictable results therefrom.
Regarding claim 9
Ma in view of Pa in view of Mi teaches or suggests:
The method of claim 7, wherein performing the search of the stored audio description data structures comprises generating a measure of similarity between characteristics of the user and characteristics of other users for which audio description data structures are stored (Ma: Col 8:49-8:58: similarities determined and used to consolidate actions of the system); (Mi: Col 14:50-14:67: system determines similar jobs of diverse users to determine and operate with respect to matching parameters of previous or historical actions to a current set of requests, steps, and output, and retrieving an audio description data structure associated with a relatively highest similarity other user as the matching audio description data structure. Examiner takes official notice that using measures of similarity to determine matches and subsequent steps based thereon was well known in the art before the effective filing date of the instant invention and would have comprised an obvious inclusion for at least the purpose of storing primitives or consistent similar tasks with respect to a user and utilizing such stored primitives to determine, process, and invoice similar users based thereon thereby providing one or more audio description augmented videos more efficiently. The claim is considered obvious over Ma as modified by Pa and Mi as addressed in the base claim as it would have been obvious to apply the further teaching of Ma, Pa, and/or Mi to the modified device of Ma, Pa, and Mi; one of ordinary skill in the art would have expected only predictable results therefrom.
Regarding claim 17—the claim is considered to recite substantially similar subject matter to that of claim 7 and is similarly rejected.
Regarding claim 18—the claim is considered to recite substantially similar subject matter to that of claim 8 and is similarly rejected.
Regarding claim 19—the claim is considered to recite substantially similar subject matter to that of claim 9 and is similarly rejected.
Conclusion
Any inquiry concerning this communication or earlier communications from the examiner should be directed to PAUL C MCCORD whose telephone number is (571)270-3701. The examiner can normally be reached 730-630 M-F.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, CAROLYN EDWARDS can be reached at (571) 270-7136. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/PAUL C MCCORD/ Primary Examiner, Art Unit 2692