DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Response to Arguments
Applicant’s arguments, see Remarks, filed 07/01/2026, with respect to the rejection of claims 21-40 under 35 U.S.C. § 101 have been fully considered and are persuasive. The rejection of claims 21-40 under 35 U.S.C. § 101 has been withdrawn.
Applicant's arguments filed 07/01/2026 have been fully considered but they are not persuasive.
Regarding the rejections under 35 U.S.C. § 103, applicant argues:
“…the Examiner admitted that Smith and Poirier fail to teach multithreaded wake-word processing…”
…
“Given this admitted deficiency of Smith and Poirier, one of ordinary skill in the art faced with Smith and Poirier would not have been reasonably led to Applicant's invention, including having a multithreaded processor generate and execute multiple wake-word processing threads in search of a wake-word signature. Furthermore, Applicant has not found in or from the newly-cited Poirier reference any disclosure or suggestion that would have reasonably led one of ordinary skill in the art to overcome what the Examiner admitted Smith fails to disclose. Poirier teaches a continuous dictation system where an audio stream is broken down into sequential, chronological segments that represent sequential spoken words, so as to efficiently transcribe an entire sentence. (See, e.g., Poirier, column 3, lines 14-23; column 8, lines 34-38.) Poirier provides no mechanism, teaching, or motivation whatsoever to distribute separate slices of an audio file to multiple separate threads for the purpose of detecting a specific wake-word signature. At a minimum, Poirier fails to teach ‘each wake-word processing thread processing a respective separate portion of the multiple separate portions of the media content item in search of sound matching a sound signature of a wake word’. In Poirier, the parallel processing is sentence- wide, not wake-word specific. Poirier discloses specifically dividing a spoken sentence into separate words (e.g., based on pauses between the words), and transcribing each separate word of the sentence simultaneously. (See, e.g., Poirier at column 6, lines 48-67.) Whereas, Applicant's 14 claims recite that the act of a multithreaded processor executing the multiple generated wake-word processing threads involves the multiple wake-word processing threads searching different respective portions of the media content item for sound matching the sound signature of the same wake-word as each other.”
Regarding applicant’s arguments, the examiner respectfully disagrees. The examiner contends that one of ordinary skill in the art faced with Smith and Poirier, would have reasonably arrived in the applicant’s invention. The examiner contends that Poirier provides Smith with the capability of using multiple single binary speech engines, such as shown in Fig. 2, to process portions of incoming audio, and use a single word vocabulary (102), such as shown in Fig. 1, where one of ordinary skill in the art may reasonably use the single word vocabulary (102) to search for a “wake-word”, as described in Smith. As disclosed in Poirier, the single binary search engines are described in Col. 3, lines 14-29:
FIG. 2 illustrates many Single Binary Speech Engines parallel processing speech audio stream. 1) The Audio Stream (210) is input into a Divider (211) where the audio stream is divided up into Audio Sub-Events (200, 202, 204, 206, 208). 2) Each Audio Sub-Event is input into an array of Single Binary Speech Engines (201, 203, 205, 207, 209). 3) If a match is true then the word is written to the Text Document (212) and that Sub-Event is removed from further processing in the speech engine array
Thus, Poirier provides the functionality of a multithreaded processor by using the single binary speech engines to process portions of an audio stream, in a parallel fashion, and match the audio to words in the single word vocabulary to create the text document. Although, the purpose of Poirier’s invention is to transcribe an audio stream, the above noted components of Poirier may be reasonably applied to Smith’s invention, as explained by the examiner, to produce the results of wake-word detection within audio media.
Additionally, the examiner contends that the applicant appears to mischaracterize the examiner’s rejection. The examiner’s rejection does not simply admit that there is a “deficiency” in the cited references of Smith and Poirier, as alleged by the applicant. The examiner simply notes that the Poirier reference does not explicitly mention that the parallel processing is applied to determine a “wake word”. That is, although Poirier determines words spoken in the incoming audio stream using a single word vocabulary, Poirier does not explicitly mention that the determined words are specifically “wake words”. However, as clearly argued by the examiner, one of ordinary skill in the art would have found it obvious to apply this functionality to Smith’s invention of false wake word identification in media content, to reasonably arrive at the claimed invention.
Furthermore, the examiner rejects applicant’s characterization of Poirier that: “the parallel processing is sentence- wide, not wake-word specific.” As noted before, Poirier provides for using a single word vocabulary to search for words within the incoming audio stream. One of ordinary skill in the art may reasonably use a single word vocabulary that contains only “wake-words” in order to arrive at the claimed invention. For these reasons, the examiner maintains that the claimed invention is obvious in view of the cited references of Smith and Poirier.
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claims 21-24, 28, 31-36, 39 and 40 are rejected under 35 U.S.C. 103 as being unpatentable over Smith (US PG Pub 20200090646) in view of Poirier et al. (US Patent 10,109,279; hereinafter “Poirier”).
As per claims 21 and 34, Smith discloses: A method and media-delivery system comprising: a memory (Smith; Fig. 2A, item 213; p. 0055 - the playback device 102 includes at least one processor 212, which may be a clock-driven computing component configured to process input data according to instructions stored in memory 213); and at least one processor, in communication with the memory, that causes the media-delivery system to carry out operations including (Smith; Fig. 2A, item 212; p. 0055 - the memory 213 may be data storage that can be loaded with software code 214 that is executable by the processor 212 to achieve certain functions): detecting, based on the executing, at least one instance of the wake word within the media content item (Smith; p. 0028-0030 - The wake-word engine suppressor may be configured to identify one or more false wake words in an audio stream that is to be output by the playback device, and when a false wake word is identified, the wake-word engine suppressor may be configured to temporarily deactivate the playback device's primary wake-word engine and cause the playback device to temporarily deactivate the wake-word engines of one or more other NMDs (e.g., one or more NMD-equipped playback devices); see also p. 0032 – keyword spotting; see also p. 0112); and whereby the detecting of the wake word facilitates having a voice-enabled device avoid responding to the wake word during playout of the media content item (Smith; p. 0028-0030 - The wake-word engine suppressor may be configured to identify one or more false wake words in an audio stream that is to be output by the playback device, and when a false wake word is identified, the wake-word engine suppressor may be configured to temporarily deactivate the playback device's primary wake-word engine and cause the playback device to temporarily deactivate the wake-word engines of one or more other NMDs (e.g., one or more NMD-equipped playback devices); see also p. 0032 – keyword spotting; see also p. 0112). Smith, however, fails to disclose generating, by a multithreaded processor, multiple wake-word processing threads; parsing a media content item into multiple separate portions; executing, by the multithreaded processor, the multiple generated wake-word processing threads, with each wake-word processing thread processing a respective separate portion of the multiple separate portions of the media content item in search of sound matching a sound signature of a wake word, thereby expediting searching for the wake word within the media content item; and generating, based on the detecting, metadata indicating time of occurrence within the media content item of the at least one instance of the wake word. Poirier does teach generating, by a multithreaded processor, multiple wake-word processing threads (Poirier; Fig. 3, items 300, 301 and 302; Col. 6, lines 48-67 - Referring to FIG. 3, there are 3 arrays of Binary Speech Engines 300, 301, and 302 (multiple processing threads)); parsing a media content item into multiple separate portions (Poirier; Fig. 3, items 304; Col. 6, lines 48-67 - The Audio Stream of spoken words enters a Divider function (304) where individual words or phrases are separated from the entire spoke audio stream); executing, by the multithreaded processor, the multiple generated wake-word processing threads, with each wake-word processing thread processing a respective separate portion of the multiple separate portions of the media content item in search of sound matching a sound signature of a wake word, thereby expediting searching for the wake word within the media content item (Poirier; Fig. 3, items 304; Col. 6, lines 48-67 - The separated words continue on to an array of Binary Speech Engines as can be seen that Word 1 (305 and 308) is provided to all the binary speech engines in Binary Speech Engine (BSE) 1 Array (300). Items 305 and 308 are the same, just shown as output of the divider stage and input as the recognition stage. In this example there are 3 Binary Speech Engine Arrays BSE 1 (300), BSE 2 (301) and BSE 3 (302). As the separate audio words are output from the divider (305, 306, and 307 in this example) they are parallel processed (multithreaded processing) in separate BSE arrays to increase processing speed through parallelism); and generating, based on the detecting, metadata indicating time of occurrence within the media content item of the at least one instance of the wake word (Poirier; Col. 8, lines 15-31 - Index the recognized word to full audio using a timestamp or some other method (708 and (709)). Although Poirier fails to explicitly teach that the threads are specifically “wake-word processing threads”, Poirier teaches keyword searching using multithreaded or parallel processing which may be combined with Smith’s wake-word processing to provide Smith’s invention with the capability of searching for a wake-word using Poirier’s multithreaded processing. Therefore, it would have been obvious to one of ordinary skill in the art to modify the method and media-delivery system of Smith to include generating, by a multithreaded processor, multiple wake-word processing threads; parsing a media content item into multiple separate portions; executing, by the multithreaded processor, the multiple generated wake-word processing threads, with each wake-word processing thread processing a respective separate portion of the multiple separate portions of the media content item in search of sound matching a sound signature of a wake word, thereby expediting searching for the wake word within the media content item; and generating, based on the detecting, metadata indicating time of occurrence within the media content item of the at least one instance of the wake word, as taught by Poirier, because parallel processing reduces transcription turnaround time (Poirier; Col. 2, lines 44-54). As per claims 22 and 35, Smith in view of Poirier discloses: The method and media-delivery system of claims 21 and 34, wherein the voice-enabled device is a voice-enabled media-playback device (Smith; p. 0022 - In turn, the VAS corresponding to the wake word that was identified by the wake-word engine receives the transmitted sound data from the NMD over a communication network. A VAS traditionally takes the form of a remote service implemented using one or more cloud servers configured to process voice inputs (e.g., AMAZON's ALEXA, APPLE's SIRI, MICROSOFT's CORTANA, GOOGLE'S ASSISTANT, etc.). In some instances, certain components and functionality of the VAS may be distributed across local and remote devices), whereby the detecting of the wake word facilitates having the voice-enabled media-playback device avoid responding to the wake word when the voice-enabled media-playback device plays the wake word within the transmitted media content item (Smith; p. 0028-0030 - The wake-word engine suppressor may be configured to identify one or more false wake words in an audio stream that is to be output by the playback device, and when a false wake word is identified, the wake-word engine suppressor may be configured to temporarily deactivate the playback device's primary wake-word engine and cause the playback device to temporarily deactivate the wake-word engines of one or more other NMDs (e.g., one or more NMD-equipped playback devices); see also p. 0032 – keyword spotting; see also p. 0112). And further, Poirier teaches transmitting to the voice-enabled media-playback device the media content item and the generated metadata (Poirier; Col. 7, lines 49-60 - Then additional audio segments that do not fall below the Sound Level threshold continued to be stored in the segment buffer. When the next audio segment threshold is below the set Sound Level trigger level the Audio Tag Marker Control (504) is checked for BOS=True. If BOS=True then this is not Begin of Sub Event (BOS) (512) and the process continues on to Tag Segment as EOS (511) and then the Audio Sub Event (i.e. the audio with a word) is created and send to the Binary Speech Engine Array to be transcribed). Although Smith fails to package the identified wake word and time of occurrence of the wake word into metadata and transmit it to the voice enabled device, Poirier provides for this functionality in the cited portions of its disclosure, so that Smith’s network of devices is able to take advantage of the transmitted metadata. Therefore, it would have been obvious to one of ordinary skill in the art to modify the method and system of Smith to include transmitting to the voice-enabled media-playback device the media content item and the generated metadata, as taught by Poirier, because parallel processing reduces transcription turnaround time (Poirier; Col. 2, lines 44-54).
As per claim 23, Smith in view of Poirier discloses: The method of claim 22, wherein transmitting the media content item to the voice-enabled media-playback device comprises streaming the media content item to the voice-enabled media-playback device (Smith; p. 0031 - In practice, the playback device may receive the audio stream that the wake-word engine suppressor analyzes via an audio interface, which may take a variety of forms and may be configured to receive audio from a variety of sources. As one example, the audio interface may take the form of an analog and/or digital line-in receptacle that physically connects the playback device to an audio source, such as a CD player or a TV).
As per claims 24 and 36, Smith in view of Poirier disclose: The method and media-delivery system of claims 22 and 35, upon which claims 24 and 36 depend. And further, Poirier teaches wherein the executing, detecting, and generating occur before transmitting the media content item to the voice-enabled media-playback device (Poirier; Col. 7, lines 49-60 - Then additional audio segments that do not fall below the Sound Level threshold continued to be stored in the segment buffer. When the next audio segment threshold is below the set Sound Level trigger level the Audio Tag Marker Control (504) is checked for BOS=True. If BOS=True then this is not Begin of Sub Event (BOS) (512) and the process continues on to Tag Segment as EOS (511) and then the Audio Sub Event (i.e. the audio with a word) is created and send to the Binary Speech Engine Array to be transcribed). Although Smith fails to package the identified wake word and time of occurrence of the wake word into metadata and transmit it to the voice enabled device, Poirier provides for this functionality in the cited portions of its disclosure, so that Smith’s network of devices is able to take advantage of the transmitted metadata. Therefore, it would have been obvious to one of ordinary skill in the art to modify the method and system of Smith to include wherein the executing, detecting, and generating occur before transmitting the media content item to the voice-enabled media-playback device, as taught by Poirier, because parallel processing reduces transcription turnaround time (Poirier; Col. 2, lines 44-54).
As per claim 28, Smith in view of Poirier disclose: The method of claim 27, wherein the metadata causes the voice-enabled media-playback device to deactivate a wake word detector of the voice-enabled media-playback device for a duration of the time range (Smith; p. 0165 - In practice, the particular amount of time for which the at least one NMD is to deactivate its wake-word engine corresponding to the identified false wake word may have a duration that is sufficient to allow the playback device 102a to output audio, via the speakers 218, that comprises the false wake word and/or to allow the NMD to receive a sound input comprising that output audio…).
As per claim 31, Smith in view of Poirier disclose: The method of claim 30, wherein the media stream comprise a live stream (Smith; p. 0118 - The voice extractor 572 transmits or streams these messages, M.sub.V, that may contain voice input in real time or near real time to a remote VAS, such as the VAS 190 (FIG. 1B), via the network interface 224).
As per claim 32, Smith in view of Poirier disclose: The method of claim 21, further comprising: determining that the media content has been updated; and responsive to determining that the media content item has been updated, repeating the executing, detecting, and generating (Smith; p. 0105 - In some embodiments, audio content sources may be added or removed from a media playback system such as the MPS 100 of FIG. 1A. In one example, an indexing of audio items may be performed whenever one or more audio content sources are added, removed, or updated. Indexing of audio items may involve scanning for identifiable audio items in all folders/directories shared over a network accessible by playback devices in the media playback system and generating or updating an audio content database comprising metadata (e.g., title, artist, album, track length, among others) and other associated information, such as a URI or URL for each identifiable audio item found. Other examples for managing and maintaining audio content sources may also be possible).
As per claim 33, Smith in view of Poirier disclose: The method of claim 21, further comprising: wherein the media-playback device plays the transmitted media content item and forwards to a voice-enabled device the metadata to cause the voice-enabled device to avoid responding to the wake word when playout of the media content item by the media-playback device includes playout of the wake word (Smith; p. 0028-0030 - The wake-word engine suppressor may be configured to identify one or more false wake words in an audio stream that is to be output by the playback device, and when a false wake word is identified, the wake-word engine suppressor may be configured to temporarily deactivate the playback device's primary wake-word engine and cause the playback device to temporarily deactivate the wake-word engines of one or more other NMDs (e.g., one or more NMD-equipped playback devices); see also p. 0032 – keyword spotting; see also p. 0112). And further, Poirier teaches transmitting the media content item and the generated metadata to a media-playback device (Poirier; Col. 7, lines 49-60 - Then additional audio segments that do not fall below the Sound Level threshold continued to be stored in the segment buffer. When the next audio segment threshold is below the set Sound Level trigger level the Audio Tag Marker Control (504) is checked for BOS=True. If BOS=True then this is not Begin of Sub Event (BOS) (512) and the process continues on to Tag Segment as EOS (511) and then the Audio Sub Event (i.e. the audio with a word) is created and send to the Binary Speech Engine Array to be transcribed). Although Smith fails to package the identified wake word and time of occurrence of the wake word into metadata and transmit it to the voice enabled device, Poirier provides for this functionality in the cited portions of its disclosure, so that Smith’s network of devices is able to take advantage of the transmitted metadata. Therefore, it would have been obvious to one of ordinary skill in the art to modify the method and system of Smith to include transmitting the media content item and the generated metadata to a media-playback device, as taught by Poirier, because parallel processing reduces transcription turnaround time (Poirier; Col. 2, lines 44-54).
As per claim 39, Smith in view of Poirier disclose: The media-delivery system of claim 34, upon which claim 39 depends. And further, Poirier teaches wherein the metadata indicates a time range that includes the wake word in the media stream (Poirier; Col. 8, lines 15-31 - Index the recognized word to full audio using a timestamp or some other method (708 and (709)). Although Poirier fails to explicitly teach that the threads are specifically “wake-word processing threads”, Poirier teaches keyword searching using multithreaded or parallel processing which may be combined with Smith’s wake-word processing to provide Smith’s invention with the capability of searching for a wake-word using Poirier’s multithreaded processing. Therefore, it would have been obvious to one of ordinary skill in the art to modify the method and media-delivery system of Smith to include wherein the metadata indicates a time range that includes the wake word in the media stream, as taught by Poirier, because parallel processing reduces transcription turnaround time (Poirier; Col. 2, lines 44-54).
As per claim 40, Smith in view of Poirier disclose: The media-delivery system of claim 39, wherein the metadata causes the voice-enabled media-playback device to deactivate a wake word detector of the voice-enabled media-playback device for a duration of the time range (Smith; p. 0028-0030 - The wake-word engine suppressor may be configured to identify one or more false wake words in an audio stream that is to be output by the playback device, and when a false wake word is identified, the wake-word engine suppressor may be configured to temporarily deactivate the playback device's primary wake-word engine and cause the playback device to temporarily deactivate the wake-word engines of one or more other NMDs (e.g., one or more NMD-equipped playback devices); see also p. 0032 – keyword spotting; see also p. 0112).
Claims 25-27, 29, 30, 37 and 38 are rejected under 35 U.S.C. 103 as being unpatentable over Smith in view of Poirier and further in view of Engineer (US PG Pub 20180160189). As per claim 25 and 37, Smith in view of Poirier disclose: The method and media-delivery system of claims 22 and 35, upon which claims 25 and 37 depend. Smith in view of Poirier, however, fail to disclose wherein the executing, detecting, and generating occur while transmitting the media content item to the voice-enabled media-playback device. Engineer does teach wherein the executing, detecting, and generating occur while transmitting the media content item to the voice-enabled media-playback device (Engineer; p. 0021 - While the video content is playing, the search component also can display the textual information (e.g., words) being spoken in the content, and can highlight or otherwise emphasize the word(s) that corresponds to (e.g., matches or substantially matches) the keyword(s) in the search as the word(s) is being displayed or scrolled across the display screen… displaying search results while media content is being streamed (transmitted)).
Therefore, it would have been obvious to one of ordinary skill in the art to modify the method of Smith and Poirier to include wherein the executing, detecting, and generating occur while transmitting the media content item to the voice-enabled media-playback device, as taught by Engineer, in order to facilitate the search and identification of keywords in media content for highlighting or emphasizing as the keywords are displayed to the user (Engineer; p. 0021).
As per claim 26, Smith in view of Poirier disclose: The method of claim 21, upon which claim 26 depends.
Smith in view of Poirier, however, fail to disclose wherein the metadata includes a first time indicating a start of a time range that includes the wake word in the media content item. And further, Engineer teaches wherein the metadata includes a first time indicating a start of a time range that includes the wake word in the media content item (Engineer; p. 0021 - The presentation of the content can start from a time position associated with a time indicator, for example, in response to a user selecting to play the content or selecting the time indicator; see also p. 0037-0038 - With regard to each word in an item of content that relates to a keyword associated with the search query, the search component 118 can generate time information and/or an indicator (e.g., a time location indicator) to facilitate indicating a time location in the item of content where the word is located; see also p. 0073). Therefore, it would have been obvious to one of ordinary skill in the art to modify the method of Smith and Poirier to include wherein the metadata includes a first time indicating a start of a time range that includes the wake word in the media content item, as taught by Engineer, in order to facilitate the search and identification of keywords in media content for highlighting or emphasizing as the keywords are displayed to the user (Engineer; p. 0021).
As per claims 27, Smith in view of Poirier and Engineer disclose:
The method of claim 26, upon which claim 27 depends. And further, Engineer teaches wherein the metadata includes a second time indicating an end of the time range that includes the wake word in the media content item (Engineer; p. 0021 - The presentation of the content can start from a time position associated with a time indicator, for example, in response to a user selecting to play the content or selecting the time indicator; see also p. 0037-0038 - With regard to each word in an item of content that relates to a keyword associated with the search query, the search component 118 can generate time information and/or an indicator (e.g., a time location indicator) to facilitate indicating a time location in the item of content where the word is located; see also p. 0073). Therefore, it would have been obvious to one of ordinary skill in the art to modify the method of Smith and Poirier to include wherein the metadata includes a second time indicating an end of the time range that includes the wake word in the media content item, as taught by Engineer, in order to facilitate the search and identification of keywords in media content for highlighting or emphasizing as the keywords are displayed to the user (Engineer; p. 0021).
As per claim 29, Smith in view of Poirier and Engineer disclose: The method of claim 27, upon which claim 29 depends. And further, Engineer wherein the first time is indicated by a first offset from a start time of the media content item, and wherein the second time is indicated by a second offset from the start time of the media content item (Engineer; p. 0021 - The presentation of the content can start from a time position associated with a time indicator, for example, in response to a user selecting to play the content or selecting the time indicator; see also p. 0037-0038 - With regard to each word in an item of content that relates to a keyword associated with the search query, the search component 118 can generate time information and/or an indicator (e.g., a time location indicator) to facilitate indicating a time location in the item of content where the word is located; see also p. 0073). Therefore, it would have been obvious to one of ordinary skill in the art to modify the method of Smith and Poirier to include wherein the first time is indicated by a first offset from a start time of the media content item, and wherein the second time is indicated by a second offset from the start time of the media content item, as taught by Engineer, in order to facilitate the search and identification of keywords in media content for highlighting or emphasizing as the keywords are displayed to the user (Engineer; p. 0021).
As per claims 30 and 38, Smith in view of Poirier disclose:
The method and media-delivery system of claims 21 and 34, wherein the media content item comprises media stream that a media-delivery system receives and forwards in real-time to the voice-enabled media-playback device (Smith; p. 0118 - The voice extractor 572 transmits or streams these messages, M.sub.V, that may contain voice input in real time or near real time to a remote VAS, such as the VAS 190 (FIG. 1B), via the network interface 224). Smith in view of Poirier, however, fail to disclose wherein the media content item comprises media stream that a media-delivery system receives and forwards in real-time to the voice-enabled media-playback device. Engineer teaches wherein the media content item comprises media stream that a media-delivery system receives and forwards in real-time to the voice-enabled media-playback device (Engineer; p. 0022 & p. 0024 - In certain implementations, the search component can be contained in, and executed in, a media device, such as, for example, a set-top box (STB) or set-top unit (STU), which can be associated with (e.g., communicatively connected to) a presentation component (e.g., a television) or another type of communication device (e.g., mobile phone, electronic pad or tablet, electronic notebook, computer, . . . ); see also p. 0082 – application component communicating with media device for streaming media content (real-time) and presentation component (for display of media content)), wherein the executing, detecting, and generating are carried out by the media-delivery system as the media-delivery system receives and forwards the media stream to the voice-enabled media-playback device (Engineer; p. 0082 – application component communicating with media device for streaming media content (real-time) and presentation component (for display of media content)). Therefore, it would have been obvious to one of ordinary skill in the art to modify the method and system of Smith and Poirier to include wherein the media content item comprises media stream that a media-delivery system receives and forwards in real-time to the voice-enabled media-playback device, wherein the executing, detecting, and generating are carried out by the media-delivery system as the media-delivery system receives and forwards the media stream to the voice-enabled media-playback device, as taught by Engineer, in order to facilitate the search and identification of keywords in media content for highlighting or emphasizing as the keywords are displayed to the user (Engineer; p. 0021).
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure. The prior art made of record and not relied upon includes: Lang (US PG Pub 20190043492) discloses example techniques involving determining a direction of a NMD. An example implementation includes a playback device receiving data representing audio content for playback by the playback device. Before the audio content is played back by the playback device, the playback device detects, in the audio content, one or more wake words for one or more voice services. The playback device causes one or more networked microphone devices to disable its respective wake response to the detected one or more wake words during playback of the audio content by the playback device and plays back the audio content via one or more speakers. When enabled, the wake response of a given networked microphone device to a particular wake word causes the given networked microphone device to listen, via a microphone, for a voice command following the particular wake word (Lang; Abstract).
THIS ACTION IS MADE FINAL. Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a).
A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to Rodrigo A Chavez whose telephone number is (571)270-0139. The examiner can normally be reached Monday - Friday 9-6 ET.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Richemond Dorvil can be reached on 5712727602. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/RODRIGO A CHAVEZ/Examiner, Art Unit 2658
/RICHEMOND DORVIL/Supervisory Patent Examiner, Art Unit 2658