Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Claim Rejections - 35 USC § 112
The following is a quotation of 35 U.S.C. 112(b):
(b) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention.
The following is a quotation of 35 U.S.C. 112 (pre-AIA ), second paragraph:
The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the applicant regards as his invention.
Claim 7 recites the limitation "the at least one annotation" in line 3 of the claim. There is insufficient antecedent basis for this limitation in the claim.
Claim 9 recites the limitation "the at least one annotation" in lines 3-. There is insufficient antecedent basis for this limitation in the claim.
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claims 1, 7, 9, 11-12, and 16-17 are rejected under 35 U.S.C. 103 as being unpatentable over Jhawar et al. (US 20190205340 A1; hereinafter Jhawar) in view of Pathak et al. (US 20190294883 A1; hereinafter Pathak).
Regarding claim 11, Jhawar teaches a computer system, comprising: one or more computer-readable storage media (“Computing device 1000 typically includes a variety of computer-readable media… By way of example, and not limitation, computer-readable media may comprise computer storage media and communication media.” (pages 9-10, para [0093]-[0094]; page 8, para [0090]; Fig 10).);
a processor set (“…computing device 1000 includes a bus 1010 that directly or indirectly couples the following devices: memory 1012, one or more processors 1014, one or more presentation components 1016...” (page 8, para [0092]; Fig 10). A processor set includes one or more processors.); and
program instructions stored on the one or more computer-readable storage media to cause the processor set to perform operations (“Computer storage media includes both volatile and nonvolatile, removable and non-removable media implemented in any method or technology for storage of information such as computer-readable instructions, data structures, program modules, or other data,” (page 9, para [0094]; page 8, para [0091]). Program instructions include computer-readable instructions.) comprising:
obtaining annotated data from an extended reality (XR) device (“Turning now to FIG. 1, an inspection environment 100 comprising a radio tower 110 being inspected by an inspector 118 who is wearing a head-mounted display that captures video of the inspection… The inspector 118 can issue audible tagging commands as the inspection progresses…,” (page 2, para [0026]-[0028]).
“Aspects of the technology described herein can be performed by a head-mounted display. The head-mounted display can include an augmented reality display,” (page 2, para [0024]).
“At step 810, an audio signal is received at a head-mounted display device while the head-mounted display device is recording a video of a scene. The audio signal captures a voice of a user of the head-mounted display device. The audio signal can be received via a microphone associated with the head-mounted display device,” (page 7, para [0077]).
Annotated data includes the captured video with audible commands / captured voice of a user. An extended reality device includes the augmented reality display.),
the annotated data comprising video data and audio data (“At step 810, an audio signal is received at a head-mounted display device while the head-mounted display device is recording a video of a scene. The audio signal captures a voice of a user of the head-mounted display device. The audio signal can be received via a microphone associated with the head-mounted display device,” (page 7, para [0077]).), wherein the annotated data is associated with a quality inspection operation (“Aspects of the technology described herein allow a user to add tags to a video as the video is being recorded. The tags can be added by capturing the user's voice as the video is recorded. A trigger word, such as “tag,” can be used to activate the tag function. For example, a user could say, “start tag, living room fireplace inspection” to insert a tag. The start method command in conjunction with the tag command can be used to tag a duration of video. In this example, upon completing a recording of the fireplace inspection the same user could say, “stop tag” to mark the end of the section,” (page 2, para [0020]).
“For example, project types can include construction inspection, sales inspection, training video, maintenance project, and others,” (page 4, para [0037]).);
detecting, based on the annotated data, at least one annotation caption indicative of an inspection subject (“At step 820, audio analysis is performed on the audio signal to identify a tag initiation command issued by a user of the head-mounted display device. The tag initiation command comprises a tag activation word and a tag description. The audio processing can be performed by the head-mounted display device. The tag description can be the name of a tag or some other way to identify a tag. For example, if a list of tags are displayed with numbers/letters delineator, then the delineator could serve as the description. For example, the user could say, “insert tag number 1” or “start tag inspection number 1.” In both examples, the “tag” can be the tag activation word,” (page 7-8, para [0077]-[0078]).
“The first entry 620 has a tag identification of “section 2 junction box 1” (i.e., section 2 of the radio tower 110), is associated with progress point 0:54, and has a unique ID of tag one. The user could associate this tag with their progress point in a video by speaking, “section 2 junction box 1.” The second entry 622 has a tag identification of “north guy wire attachment,” a progress point of 1:12 and a unique ID of tag 2. The user could associate this tag with their progress point in a video by speaking, “north guy wire attachment,” (page 7, para [0065]; page 6, para [0062]; Fig 6).
“The audio processing component 216 processes an audio signal and recognizes human speech within the audio signal. The human speech can be converted to text by the audio processing component. The text can then be used to control tagging functions…,” (page 4, para [0042]).
Detecting at least one annotation caption includes identifying a tag initiation command. Detecting includes the audio analysis / identification. At least one annotation caption includes a tag initiation command and / or a tag description. The detecting is based on the captured video with audible commands / captured voice of a user (annotated data).);
determining at least one annotation interval based on the at least one annotation caption, the at least one annotation interval comprising at least one of a video interval or an audio interval (“a user could say, ‘start tag, living room fireplace inspection’ to insert a tag. The start method command in conjunction with the tag command can be used to tag a duration of video. In this example, upon completing a recording of the fireplace inspection the same user could say, ‘stop tag’ to mark the end of the section. An alternative tag method is a point-in-time tag that could be created by the command ‘insert tag living room maintenance.’ The point in time tag is associated with a single progress point in the video recording,” (page 2, para [0020]; page 8, para [0079]).
“…tags can be for a point in time or duration of video. The duration tagging can use a start and end command,” (page 5, para [0049]).
At least one annotation caption includes a start method command and / or stop command. At least one annotation interval includes the duration of video. The duration of a video is determined based on the start and stop / end commands.);
generating tagged data by tagging the at least one of the video interval or the audio interval based on the annotation caption (“At step 830, an association of the tag description with a particular duration point of the video is stored in a computer storage, as described with reference to FIG. 6. The particular duration point coincides approximately with a point in time when the audio signal was received. In one aspect, an image of the scene at the particular duration point is also captured and stored. The image can be associated with the tag,” (page 8, para [0081]; page 7, para [0075]; Fig 6; Fig 7; Fig 8).
“In one aspect, a tag is associated with a progress point, such as 0:54 610 or 1:12 612... The tag can comprise metadata stored in a separate file that is linked to the progress point. Here two tag entries are shown. The first entry 620 has a tag identification of ‘section 2 junction box 1’ (i.e., section 2 of the radio tower 110), is associated with progress point 0:54, and has a unique ID of tag one. The user could associate this tag with their progress point in a video by speaking, ‘section 2 junction box 1.’” (page 7, para [0065]-[0066]; Fig 6).
Tagging the at least one of the video interval includes associating a tag description / identification with a particular duration / progress point of the video. Tagging the at least one of the video interval is based on the tag description. Tagged data includes the first entry 620.); and
transmitting the tagged data to a data store (“The user devices can send and receive communications, including video and images, project information, tag collections, and tags,” (page 3, para [0030]-[0033]).
“The project component 218 can also associate videos generated by the HMD with the project for subsequent retrieval by the HMD, a different HMD, or another computing device. The project component 218 can store tagged videos and retrieve tagged videos based on a project designation,” (page 4-5, para [0044]; Fig 2).
A data store includes project component 218. Transmitting the tagged data includes sending, receiving and / or retrieving tagged videos.) to facilitate presentation, by an output component of a computing device, a representation of the tagged data (“…videos capturing a construction inspection or other operation may be tagged for navigational purposes. The tagging may need to follow a specific schema, rather than a freestyle approach, in order to facilitate subsequent navigation of a recorded video or satisfy the requirements of an inspection project,” (page 5, para [0045]).
“Tags may be used for a number of purposes, such as… points to snap the video to upon selection of a tag,” (page 1, para [0002]).
“Presentation component(s) 1016 presents data indications to a user or other device. Exemplary presentation components include a display device, speaker, printing component, vibrating component, and the like,” (page 9, para [0096]; page 8, para [0092]).
Facilitating presentation includes facilitating subsequent navigation of a recorded video. An output component of a computing device includes a presentation component.).
Jhawar is not relied upon teaching but Pathak teaches the at least one annotation caption is indicative of a defect (“The machine learning component 112 can also display additional information overlay 906 that classifies the type of part defect as having missing coating. In this example, the additional information overlay 906 is a box that highlights the area of the defect and classifies the defect as having missing coating. The annotation can be by a box or by any other shape or even by highlighting the zone of defect in a different color,” (page 5, para [0048]).
“The machine learning component 112 can also automatically tag and overlay information with the video frames, which can be displayed as an augmented reality layer on the raw video feed. The overlay information can be, but is not limited to, the type of part defect, location of defect, size of defect, engine history, operational characteristics, etc,” (page 3, para [0033]).
“The machine learning component 112 can classify the type of part defect and determine the location of defect. The type of part defect and location of defect can be, for example, based on historical data from past videos (e.g., previous database of images)… For example, the machine learning component 112 can employ deep learning and feature recognition for distress or failure level identification,” (page 3, para [0032]).
“…an inspector can provide input and modify the classification of the type of part defect in the overlay information,” (page 4, para [0035]).
After combination, Pathak’s overlay information becomes Jhawar’s tag description (annotation caption). Pathak’s overlay information is indicative of a defect.)
Before the effective filling date of the claimed invention, it would have been obvious to one having ordinary skill in the art to apply the teachings of Pathak to Jhawar. The motivation would have been to “pull up information of a defect more easily,” (Pathak; page 3, para [0031]). Additional motivation would have been to “speed up the process of… inspection, [and] reduce misses of defect due to oversight,” (Pathak; page 2, para [0020]). Additional motivation would have been to save time reviewing inspection video.
Regarding claims 1 and 17, they are rejected using the same citations and rationales described in the rejection of claim 11.
Regarding claim 12, Jhawar in view of Pathak teaches the computer system of claim 11, wherein detecting the at least one annotation caption comprises: determining a defect type comprising at least one of an audio defect or a video defect (Pathak; “a selection component that analyzes the video feed and identifies frames that capture information of part damage and defects. The computer executable components can further comprise a machine learning component that classifies type of part defect, determines location of defect and learns the digital grid,” (page 1, para [0003]).
“The machine learning component 112 can classify the type of part defect and determine the location of defect,” (Pathak; page 3-4, para [0032]-[0034]).
“The machine learning component 112 can employ deep learning and feature recognition using previous database of images to identify distress or failure level (e.g., defect). The machine learning component 112 can display the key features along with the classification of the type of part defect and the location of the defect as an augmented reality layer on the raw video feed,” (Pathak; page 5, para [0047]-[0048]; Fig 9).
Determining a defect type includes classifying the type of part defect. Comprising a video defect includes the defect is identified from the video frames / data.).
Before the effective filling date of the claimed invention, it would have been obvious to one having ordinary skill in the art to apply the teachings of Pathak to Jhawar. The motivation would have been to “pull up information of a defect more easily,” (Pathak; page 3, para [0031]). Additional motivation would have been to “speed up the process of… inspection, [and] reduce misses of defect due to oversight,” (Pathak; page 2, para [0020]). Additional motivation would have been to save time reviewing inspection video.
Regarding claim 7, Jhawar in view of Pathak teaches the computer-implemented method of claim 1, determining the at least one annotation interval comprises segmenting the video data to isolate a frame corresponding to a defect location identified by the at least one annotation (Pathak; “The selection component 110 can select from a sequence of static frames (not shown) the image 904 as the frame that captures the most information (e.g., best image)… The machine learning component 112 can also display additional information overlay 906 that classifies the type of part defect as having missing coating. In this example, the additional information overlay 906 is a box that highlights the area of the defect and classifies the defect as having missing coating. The annotation can be by a box or by any other shape or even by highlighting the zone of defect in a different color.” (Pathak; page 5, para [0048]; Fig 9).
“the computer-implemented method 500 can comprise analyzing (e.g., via the selection component 110), by the system, the video feed and identifying frames that capture information of part damage and defects. The selection component 110 can employ video abstraction techniques to extract informative frames from the borescope inspection videos. The selection component 110 can analyze and select from a sequence of frames the frame that are informative as to the type of defect, location of defect, etc,” (Pathak; pages 4-5, para [0042]).
“The machine learning component 112 can also learn the type and location of defect from input data from an inspector,” (Pathak; page 4, para [0034]).
Segmenting the video data to isolate a frame includes selecting a frame that captures the most information (e.g., best image). The selected frame corresponds to a box (annotation) that highlights the area of the defect.).
Before the effective filling date of the claimed invention, it would have been obvious to one having ordinary skill in the art to apply the teachings of Pathak to Jhawar. The motivation would have been to “pull up information of a defect more easily,” (Pathak; page 3, para [0031]). Additional motivation would have been to “speed up the process of… inspection, [and] reduce misses of defect due to oversight,” (Pathak; page 2, para [0020]). Additional motivation would have been to save time reviewing inspection video.
Regarding claim 9, Jhawar in view of Pathak teaches the computer-implemented method of claim 1, wherein tagging the at least one of the video interval or the audio interval comprises associating at least one tag with the at least one annotation interval (Jhawar; “Turning now to FIG. 6, a tagged video 600 is illustrated. The video comprises frames 601, 602, 603, 604, and 605. In one aspect, a tag is associated with a progress point, such as 0:54 610 or 1:12 612. Each video frame may be uniquely identified, for example, each frame could have a number or progress point. The tag can comprise metadata stored in a separate file that is linked to the progress point. Here two tag entries are shown. The first entry 620 has a tag identification of “section 2 junction box 1” (i.e., section 2 of the radio tower 110), is associated with progress point 0:54, and has a unique ID of tag one,” (page 7, para [0065]; page 8, para [0081]; Fig 6).
“‘Start’ is one example of a tagging method command that can be part of the tag initiation command. The ‘start’ method command can initiate tagging a length or duration of video. The start command and can be paired with a stop command to stop the tagging process, such as, ‘stop tag’,” (Jhawar; page 8, para [0079]).
“…tags can be for a point in time or duration of video. The duration tagging can use a start and end command,” (Jhawar; page 5, para [0049]).
“For example, a user could say, “start tag, living room fireplace inspection” to insert a tag. The start method command in conjunction with the tag command can be used to tag a duration of video. In this example, upon completing a recording of the fireplace inspection the same user could say, “stop tag” to mark the end of the section,” (Jhawar; page 2, para [0020]).
At least one annotation interval includes the duration of video. Associating at least one tag with the at least one annotation interval includes associating a tag with a progress point of the video / duration of video.),
the tag being indicative of a defect characteristic identified from the at least one annotation (Pathak; “The machine learning component 112 can also automatically tag and overlay information with the video frames, which can be displayed as an augmented reality layer on the raw video feed. The overlay information can be, but is not limited to, the type of part defect, location of defect, size of defect, engine history, operational characteristics, etc,” (Pathak; page 3, para [0033]).
“The machine learning component 112 can also display additional information overlay 906 that classifies the type of part defect as having missing coating. In this example, the additional information overlay 906 is a box that highlights the area of the defect and classifies the defect as having missing coating. The annotation can be by a box or by any other shape or even by highlighting the zone of defect in a different color,” (Pathak; page 5, para [0048]).
“…an inspector can provide input and modify the classification of the type of part defect in the overlay information,” (Pathak; page 3-4, para [0034]-[0035]).
Jhawar; “At step 820, audio analysis is performed on the audio signal to identify a tag initiation command issued by a user of the head-mounted display device. The tag initiation command comprises a tag activation word and a tag description. The audio processing can be performed by the head-mounted display device. The tag description can be the name of a tag or some other way to identify a tag. For example, if a list of tags are displayed with numbers/letters delineator, then the delineator could serve as the description. For example, the user could say, “insert tag number 1” or “start tag inspection number 1.” In both examples, the “tag” can be the tag activation word,” (page 7-8, para [0077]-[0078]; Jhawar; page 5, para [0048]).
After combination, Pathak’s overlay information becomes Jhawar’s tag description (annotation caption). Pathak’s overlay information is indicative of a defect. A defect characteristic includes the type of part defect, location of defect, size of defect. After combination, Jhawar’s tag initiation command (annotation caption) includes Pathak’s inspector input. Identifying from the at least one annotation includes identifying a tag initiation command issued by a user. ).
Before the effective filling date of the claimed invention, it would have been obvious to one having ordinary skill in the art to apply the teachings of Pathak to Jhawar. The motivation would have been to “pull up information of a defect more easily,” (Pathak; page 3, para [0031]). Additional motivation would have been to “speed up the process of… inspection, [and] reduce misses of defect due to oversight,” (Pathak; page 2, para [0020]). Additional motivation would have been to save time reviewing inspection video.
Regarding claim 16, Jhawar in view of Pathak teaches the computer system of claim 11, wherein the tagged data comprises metadata including at least one of a timestamp, a defect type identifier, or text associated with the defect (Jhawar; “The tag can comprise metadata stored in a separate file that is linked to the progress point. Here two tag entries are shown. The first entry 620 has a tag identification of “section 2 junction box 1” (i.e., section 2 of the radio tower 110), is associated with progress point 0:54, and has a unique ID of tag one,” (page 7, para [0065]; Fig 6).
“The tags are time coded according to a progress point in a video recording. A video progress point is measured from the starting point to a point in the video in units of time, such as seconds. Accordingly, a video that has been recording for 45 minutes and 30 seconds has a progress point of 45 minutes and 30 seconds,” (Jhawar; page 2, para [0021]).
A timestamp includes a progress point.).
Claim 2 is rejected under 35 U.S.C. 103 as being unpatentable over Jhawar in view of Pathak in further view of Candelore (US 20240194205 A1; hereinafter Candelore).
Regarding claim 2, Jhawar in view of Pathak teaches the computer-implemented method of claim 1, wherein determining the at least one annotation caption comprises detecting at least one user command based on the audio data and a audio processing component (Jhawar; “At step 820, audio analysis is performed on the audio signal to identify a tag initiation command issued by a user of the head-mounted display device. The tag initiation command comprises a tag activation word and a tag description… For example, if a list of tags are displayed with numbers/letters delineator, then the delineator could serve as the description. For example, the user could say, ‘insert tag number 1’ or ‘start tag inspection number 1.’ In both examples, the ‘tag’ can be the tag activation word,” (pages 7-8, para [0078]-[0079]).
“The audio processing component 216 processes an audio signal and recognizes human speech within the audio signal. The human speech can be converted to text by the audio processing component. The text can then be used to control tagging functions,” (Jhawar; page 4, para [0042]).
“The audio signal can be processed with an acoustic model, which identifies sounds within the audio signal. The sounds can then be processed by a language model, which matches the sounds to words, phrases, and or sentences,” (Jhawar; page 4, para [0043]).
At least one annotation caption includes a tag initiation command. Detecting at least one user command includes recognizing human speech such as when a user says “insert tag number 1” or “start tag inspection number 1.” The detecting is based on the audio signal (audio data) and the audio processing component. The audio processing component comprises an acoustic model and a language model.).
Jhawar in view of Pathak is not relied upon teaching the detecting at least one user command is based on a natural language processing (NLP) model.
Candelore teaches detecting at least one user command is based on a natural language processing (NLP) model (“In accordance with an embodiment, the wearable device 102 may detect, in the captured one or more second audio signals, the voice command. The voice command may be detected based on an application of the natural language processing model (such as, the NLP model 212 of FIG. 2) on the one or more captured second audio signals. The wearable device 102 may determine one or more instructions specified by the user 112 through the voice command, based on, for example, the natural language processing model (such as, the NLP model 212 of FIG. 2),” (page 4, para [0033]).).
Before the effective filling date of the claimed invention, it would have been obvious to one having ordinary skill in the art to apply the teachings of Candelore to Jhawar in view of Pathak. The motivation would have been to improve semantic understanding of audio signals / data. Additional motivation would have been to improve the recognition of the intent and / or meaning behind user commands / speech. Additional motivation would have been to improve the handling and / or understanding of speech variations.
Claims 3 and 19 are rejected under 35 U.S.C. 103 as being unpatentable over Jhawar in view of Pathak in further view of Shroff et al. (US 20250030940 A1; hereinafter Shroff).
Regarding claim 19, Jhawar in view of Pathak is not relied upon teaching but Shroff teaches the computer program product of claim 17, wherein detecting the at least one annotation caption comprises detecting a user gesture in the video data using a trained gesture recognition model, the user gesture comprising a gesture by a hand of the user that indicates a defect location (“Step 502 includes receiving, in a first camera device mounted on a smart glass system, a command from a user, the command being indicative of an area of interest in a forward-image viewed by the user. In some embodiments, step 502 includes identifying the command as a hand gesture from the user with a hand gesture recognition model based on a learned history of recorded hand gestures from the user. In some embodiments, step 502 includes receiving one of: a two-hand gesture marking opposite corners of the area of interest, a finger gesture delineating a border of the area of interest, a two-finger gesture forming a cross-hair indicating a center of the area of interest, or a circle gesture including the center of the area of interest,” (page 3, para [0030]).
“the camera captures the scene, including some example hand gestures from the user. The frames including hand gestures are run through a gesture recognition system to recognize the moment when the appropriate gesture is presented by the user, and identify the region of interest based on a reading of the hand gesture from the user,” (page 2, para [0021]).
“Dataset 103-1 may include a recorded video, audio, or some other file or streaming media,” (page 2, para [0022]; page 1, para [0002]).
“Object of interest 229 may be identified from a scene 210 by the user via the voice command or pointed at using a hand gesture,” (page 3, para [0026]).
Detecting a user gesture includes identifying the command as a hand gesture from the user. A trained gesture recognition model includes a hand gesture recognition model based on a learned history of recorded hand gestures from the user. The user gesture comprising a gesture by a hand of the user that indicates a defect location includes a circle gesture including the center of the area of interest. After combination, Shroff’s area of interest becomes Pathak’s location of defect.).
Before the effective filling date of the claimed invention, it would have been obvious to one having ordinary skill in the art to apply the teachings of Shroff to Jhawar in view of Pathak. The motivation would have been to improve the precision and / or accuracy of user commands. Additional motivation would have been to enable user commands in loud places or when there is excessive background noise. Additional motivation would have been to enable faster actions and / or a more natural feel for the user. Additional motivation would have been to improve the user experience.
Regarding claim 3, Jhawar in view of Pathak is not relied upon teaching but Shroff teaches the computer-implemented method of claim 1, wherein determining the at least one annotation caption comprises detecting, based on the video data and a trained gesture recognition model, a gesture of a user (“Step 502 includes receiving, in a first camera device mounted on a smart glass system, a command from a user, the command being indicative of an area of interest in a forward-image viewed by the user. In some embodiments, step 502 includes identifying the command as a hand gesture from the user with a hand gesture recognition model based on a learned history of recorded hand gestures from the user. In some embodiments, step 502 includes receiving one of: a two-hand gesture marking opposite corners of the area of interest, a finger gesture delineating a border of the area of interest, a two-finger gesture forming a cross-hair indicating a center of the area of interest, or a circle gesture including the center of the area of interest,” (page 3, para [0030]).
“the camera captures the scene, including some example hand gestures from the user. The frames including hand gestures are run through a gesture recognition system to recognize the moment when the appropriate gesture is presented by the user, and identify the region of interest based on a reading of the hand gesture from the user,” (page 2, para [0021]).
“Dataset 103-1 may include a recorded video, audio, or some other file or streaming media,” (page 2, para [0022]; page 1, para [0002]).
“Object of interest 229 may be identified from a scene 210 by the user via the voice command or pointed at using a hand gesture,” (page 3, para [0026]).
Detecting a gesture of a user includes identifying the command as a hand gesture from the user. A trained gesture recognition model includes a hand gesture recognition model based on a learned history of recorded hand gestures from the user. The detecting is based on the hand gesture recognition model and the frames.).
Before the effective filling date of the claimed invention, it would have been obvious to one having ordinary skill in the art to apply the teachings of Shroff to Jhawar in view of Pathak. The motivation would have been to improve the precision and / or accuracy of user commands. Additional motivation would have been to enable user commands in loud places or when there is excessive background noise. Additional motivation would have been to enable faster actions and / or a more natural feel for the user. Additional motivation would have been to improve the user experience.
Claims 4 and 20 are rejected under 35 U.S.C. 103 as being unpatentable over Jhawar in view of Pathak in further view of Shroff in further view of Kumar et al. (US 20160091976 A1; hereinafter Kumar).
Regarding claim 4, Jhawar in view of Pathak in further view of Shroff is not relied upon teaching but Kumar teaches the computer-implemented method of claim 3, wherein detecting the gesture comprises: identifying a location of a fingertip of the user in the video data (Kumar; “a method that captures the ego-centric video containing the dynamic hand gesture, analyzes a frame of the ego-centric video to detect pixels that correspond to a fingertip using a hand segmentation algorithm, analyzes temporally one or more frames of the ego-centric video to compute a path of the fingertip in the dynamic hand gesture, localizes the region of interest based on the path of the fingertip in the dynamic hand gesture and performs an action based on an object in the region of interest,” (page 1, para [0006]; page 2, para [0026]).
“In one embodiment, a fingertip used to indicate the location of the region of interest may be detected from the binary mask 300 obtained from the hand segmentation algorithm. For example, the binary mask 300 and the hand segmentation algorithm may be used to identify one or more pixels that correspond to a hand region 306. The pixels in the hand region 306 may then be analyzed to identify the pixels with the largest or smallest coordinates along a given dimension (e.g., horizontal, vertical or diagonal, or equivalently, across columns or rows), or to identify pixels that have the most extreme coordinate values,” (pages 2-3, para [0032]).
“FIG. 4 illustrates a fingertip mark 402 that may identify a location of the fingertip 206,” (page 3, para [0033]).); and
tracking a movement of the fingertip (Kumar; “a method that captures the ego-centric video containing the dynamic hand gesture, analyzes a frame of the ego-centric video to detect pixels that correspond to a fingertip using a hand segmentation algorithm, analyzes temporally one or more frames of the ego-centric video to compute a path of the fingertip in the dynamic hand gesture, localizes the region of interest based on the path of the fingertip in the dynamic hand gesture and performs an action based on an object in the region of interest,” (page 1, para [0006]; page 2, para [0026]).
“…a line or a path 404 may be used to track or trace a path of the fingertip 206 during the course of the ego-centric video that is captured,” (page 3, para [0033]; Fig 4).
“…the dynamic hand gesture may be defined as a gesture generated by moving a hand or fingers of the hand. In other words, the hand may be located in different portions of consecutive frames of a video image, unlike a static hand gesture where a user's hand remains relatively stationary throughout each frame of the video image,” (page 2, para [0028]).).
Before the effective filling date of the claimed invention, it would have been obvious to one having ordinary skill in the art to apply the teachings of Kumar to Jhawar in view of Pathak in further view of Shroff. The motivation would have been to improve the precision and / or accuracy of user commands. Additional motivation would have been to improve spatial precision. Additional motivation would have been to enable faster actions and / or a more natural feel for the user. Additional motivation would have been to improve the user experience. Additional motivation would have been to enable the user to select regions of interest. Additional motivation would have been to enable more flexible selection of regions of interest; for example: “Methods for dynamic hand-gesture-based region of interest localization are desirable because they are not limited to specific shapes or enclosures,” (Kumar; page 1, para [0005]).
Regarding claim 20, Jhawar in view of Pathak in further view of Shroff is not relied upon teaching but Kumar teaches the computer program product of claim 19, wherein generating the tagged data comprises: identifying a path by tracking a movement of the hand in the video data (Kumar; “a method that captures the ego-centric video containing the dynamic hand gesture, analyzes a frame of the ego-centric video to detect pixels that correspond to a fingertip using a hand segmentation algorithm, analyzes temporally one or more frames of the ego-centric video to compute a path of the fingertip in the dynamic hand gesture, localizes the region of interest based on the path of the fingertip in the dynamic hand gesture and performs an action based on an object in the region of interest,” (page 1, para [0006]; page 2, para [0026]).
“…a line or a path 404 may be used to track or trace a path of the fingertip 206 during the course of the ego-centric video that is captured,” (Kumar; page 3, para [0033]; Fig 4).
“…the dynamic hand gesture may be defined as a gesture generated by moving a hand or fingers of the hand. In other words, the hand may be located in different portions of consecutive frames of a video image, unlike a static hand gesture where a user's hand remains relatively stationary throughout each frame of the video image,” (Kumar; page 2, para [0028]).
Identifying a path by tracking a movement of the hand includes computing / tracking / tracing a path of the fingertip.);
drawing the path on a frame of the video data (Kumar; “A visible marker or line may be used to display the path of the fingertip to the user on the display of the head-mounted video device,” (page 2, para [0020]).
“As illustrated in FIG. 5 beginning with the image 502, the image 502 illustrates a fingertip 206 with a location of the fingertip 206 indicated by the fingertip marker 402. The user is interested in performing a dynamic hand gesture to select a region of interest around the object 520 in the image 502. As the user moves his or her fingertip 206 around the object 520 from the image 502 to the image 512, a marker or line 404 traces a path of the fingertip 206,” (Kumar; page 3, para [0037]; Fig 5).); selecting a representative image frame from the video data, wherein the representative image frame is selected as a frame without a presence of the hand (Pathak; “The grid component 108 can employ the grid 902 to capture the image 904. The selection component 110 can select from a sequence of static frames (not shown) the image 904 as the frame that captures the most information (e.g., best image). The machine learning component 112 can automatically analyze and display a list of key features (not shown). The machine learning component 112 can also display additional information overlay 906 that classifies the type of part defect as having missing coating. In this example, the additional information overlay 906 is a box that highlights the area of the defect and classifies the defect as having missing coating. The annotation can be by a box or by any other shape or even by highlighting the zone of defect in a different color,” (page 5, para [0048]; Fig 9).
“the selection component 110 can compare the contrast levels of a sequence of static images to determine which image provide the most information,” (Pathak; page 3, para [0031])
A representative image includes the image 904 as the frame that captures the most information (e.g., best image). It is clear from Figure 9 that the best image does not contain the presence of a hand.); and drawing, on the representative image frame, a bounding box around the defect location based on the path (Pathak; “The grid component 108 can employ the grid 902 to capture the image 904. The selection component 110 can select from a sequence of static frames (not shown) the image 904 as the frame that captures the most information (e.g., best image). The machine learning component 112 can automatically analyze and display a list of key features (not shown). The machine learning component 112 can also display additional information overlay 906 that classifies the type of part defect as having missing coating. In this example, the additional information overlay 906 is a box that highlights the area of the defect and classifies the defect as having missing coating. The annotation can be by a box or by any other shape or even by highlighting the zone of defect in a different color,” (page 5, para [0048]; Fig 9).
From at least Pathak’s Figure 9, it is clear that the box is drawn on the representative frame and around the defect location.
Kumar; “the temporal hand gesture recognition module may detect specific temporal hand and finger movements tracing a path in the vicinity of a region of interest. In one embodiment, the region of interest extraction module may compute a tightest bounding box enclosing the traced path,” (Kumar; page 2, para [0026]).
“In one embodiment a shape may be fitted around the region of interest. In one embodiment, the shape may be a circle, a rectangle, a square, a polygon, and the like. In one embodiment, the shape may then be presented as an overlay onto the displayed image,” (Kumar; page 5, para [0070]).
After combination, Pathak’s box becomes a bounding box. Kumar teaches the bounding box is based on the path (includes computing a tightest bounding box enclosing the traced path).).
Before the effective filling date of the claimed invention, it would have been obvious to one having ordinary skill in the art to apply the teachings of Pathak to Jhawar. The motivation would have been to “pull up information of a defect more easily,” (Pathak; page 3, para [0031]). Additional motivation would have been to “speed up the process of… inspection, [and] reduce misses of defect due to oversight,” (Pathak; page 2, para [0020]). Additional motivation would have been to save time reviewing inspection video.
Before the effective filling date of the claimed invention, it would have been obvious to one having ordinary skill in the art to apply the teachings of Kumar to Jhawar in view of Pathak in further view of Shroff. The motivation would have been to improve the precision and / or accuracy of user commands. Additional motivation would have been to improve spatial precision. Additional motivation would have been to enable faster actions and / or a more natural feel for the user. Additional motivation would have been to improve the user experience. Additional motivation would have been to enable the user to select regions of interest / defects. Additional motivation would have been to enable more flexible selection of regions of interest; for example: “Methods for dynamic hand-gesture-based region of interest localization are desirable because they are not limited to specific shapes or enclosures,” (Kumar; page 1, para [0005]).
Claim 8 is rejected under 35 U.S.C. 103 as being unpatentable over Jhawar in view of Pathak in further view of Bufi (US 20220366558 A1; hereinafter Bufi).
Regarding claim 8, Jhawar in view of Pathak teaches the computer-implemented method of claim 7, wherein segmenting the video data further comprises generating a Pathak; “The selection component 110 can select from a sequence of static frames (not shown) the image 904 as the frame that captures the most information (e.g., best image)… The machine learning component 112 can also display additional information overlay 906 that classifies the type of part defect as having missing coating. In this example, the additional information overlay 906 is a box that highlights the area of the defect and classifies the defect as having missing coating. The annotation can be by a box or by any other shape or even by highlighting the zone of defect in a different color.” (Pathak; page 5, para [0048]; Fig 9).
“One or more embodiments described herein can… determine the location of defect,” (Pathak; page 2, para [0021]).
From at least Figure 9, it is clear that the box is generated around a defect location.).
Jhawar in view of Pathak is not relied upon teaching but Bufi teaches the box is a bounding box (“Passing the image through the network 156 generates bounding boxes with classes and confidence scores. The bounding box encloses an object (defect) located in the image. The class corresponds to a defect type,” (page 13, para [0302]).
“The defect location is transmitted to the PLC 146 with bounding box coordinates x0,y0,x1,y1. The PLC 146 can use the defect location information to pinpoint where on that specific section (e.g. specific lobe or journal of a camshaft being inspected) the defect was found,” (page 14, para [0312]).
“The defects are enclosed by bounding boxes 714, 716, and 718. The defects include a first paint defect contained in bounding box 714, a second paint defect contained in bounding box 716, and a porosity defect contained in bounding box 718,” (page 14, para [0320]; Fig 7).).
Before the effective filling date of the claimed invention, it would have been obvious to one having ordinary skill in the art to apply the teachings of Bufi to Jhawar in view of Pathak. The motivation would have been to enable more precise location of defect. Additional motivation would have been to provide coordinates of defect location. Additional motivation would have been to measure / estimate the size of a defect. Additional motivation would have been to enable tracking a defect across frames.
Claim 14 is rejected under 35 U.S.C. 103 as being unpatentable over Jhawar in view of Pathak in further view of Gerber et al. (US 11546507 B1; hereinafter Gerber).
Regarding claim 14, Jhawar in view of Pathak is not relied upon teaching but Gerber teaches the computer system of claim 11, further comprising extracting at least one annotation tag from the at least one annotation interval (“…audio file 704 has been segmented into multiple portions of audio, also referred to as ‘audio clips.’ In this case, audio file 704 has been segmented into a first audio clip 760, a second audio clip 762, a third audio clip 764, and a fourth audio clip 766. Audio files can be segmented based on fixed intervals of time, or using contextual information within the audio file. Moreover, in some cases, audio files may be segmented so that the audio clips are in one-to-one correspondence with the video clips, as in FIG. 7,” (col 9, lines 60-67; col 10, lines 1-23; Fig 7).
“the application can identify and classify audio features from audio information contained within audio file 704. In particular, the application has identified a first audio feature 720 (the audible word ‘refrigerator’) and a second audio feature 722 (the audible phrase ‘crack in window’),” (col 10, lines 24- 33).
“the lease features and/or audio features, which may be considered as tags for the lease and audio file, respectively, can be used to tag appropriate video features identified by the application. In particular, in some cases, the application may only tag and store video clips for later retrieval when the classified video feature corresponds with one or more of the lease tags and/or audio tags,” (col 10, lines 34-47).
“Based on these matches, the application may tag each video clip using the corresponding lease and/or audio tags, and then store the tagged video clips for later retrieval,” (col 10, lines 48-67; col 11 lines 1-4).
An annotation tag includes audio features, which may be considered as tags. After combination, Gerber’s video and / or audio clips become Jhawar’s duration of video / annotation intervals. From at least Figure 7, it is clear that the audio features are extracted from the video and / or audio clips.).
Before the effective filling date of the claimed invention, it would have been obvious to one having ordinary skill in the art to apply the teachings of Gerber to Jhawar in view of Pathak. The motivation would have been to “[reduce] the chance that the system will misidentify features from the video, or fail to detect and catalogue important features… explicitly mentioned by the user,” (Gerber; col 3, lines 25-44). Additional motivation would have been to improve the accuracy and / or descriptiveness of tags. Additional motivation would have been to “ensure only video clips that show features relevant for… features explicitly mentioned by the user are tagged and stored for later retrieval…,” (col 10, lines 34-47).
Claim 15 is rejected under 35 U.S.C. 103 as being unpatentable over Jhawar in view of Pathak in further view of Gerber in further view of Plantinga et al. (US 20260017314 A1; hereinafter Plantinga).
Regarding claim 15, Jhawar in view of Pathak in further view of Gerber is not relied upon teaching but Plantinga teaches the computer system of claim 14, wherein extracting the at least one annotation tag comprises extracting the at least one annotation tag using a large language model (LLM) (“The automated platform 102 may provide the pre-existing tags and the summaries 258 along with a prompt or instruction to the LLM 104 to generate relevant tags 260… This step may leverage the LLM's ability to recognize patterns, identify relationships between entities, and make predictions to create a set of relevant and meaningful tags 260 that (e.g., accurately) represent the video content... In these embodiments, the tags 260 may (e.g., only) include those that the LLM 104 generated. In some embodiments, the automated platform 102 may cause the LLM 104 to generate tags 260 directly from the transcripts 256…,” (page 4, para [0033]).
“the automated platform 102 may convert videos… into audio format 254. The conversion may be performed using a system, program, or tool (253) that is capable of processing video, audio, and/or other multimedia files or streams, and more particularly extracting audio from videos. The automated platform 102 may process the converted audio 254 to generate transcripts 256. This can involve the use of one or more natural language processing (NLP) techniques 255, such as speech-to-text (S2T) algorithm(s) (or automatic speech recognition (ASR) algorithm(s)),” (pages 3-4, para [0032]).).
Before the effective filling date of the claimed invention, it would have been obvious to one having ordinary skill in the art to apply the teachings of Plantinga to Jhawar in view of Pathak in further view of Gerber. The motivation would have been to improve the accuracy / relevancy / meaningfulness of tags. Additional motivation would have been to leverage sematic reasoning / understanding to assign tags even when exact keyword matches are missing. Additional motivation would have been to enable the handling of unstructured data. Additional motivation would have been to improve the flexibility of classifications related to tagging.
Claim 10 is rejected under 35 U.S.C. 103 as being unpatentable over Jhawar in view of Pathak in further view of Plantinga.
Regarding claim 10, Jhawar in view of Pathak teaches the computer-implemented method of claim 9, tagging the at least one of the video interval or the audio interval further comprises correcting the at least one tag by comparing the at least one tag with a predefined tag database (Jhawar; “In one aspect, only curated tags may be applied to the video. In this implementation, disambiguation of a received tag may be required. For example, the user may not precisely recite the tag identification language. In this situation, the closest tags can be retrieved in the user asked to select one of the suggested tags. The disambiguation interface could also allow the user to request the most relevant tags based on a present context,” (page 8, para [0080]; (Jhawar; page 5, para [0046]-[0048]).).
Correcting the at least one tag includes disambiguation of a received tag. Comparing the at least one tag with a predefined tag database includes correcting against curated tags.) based on a frequency or context (Jhawar; “The starting point can be a list of available tags. In one aspect, the available tags are ranked by likelihood of use. In one aspect, the most frequently used tags are displayed. The most frequently used tags can be determined by analyzing tag data for previously tagged videos. Tags from videos that share characteristics (e.g., location, project, user, company, venue) can be given more weight when determining the most likely tags to be used. In one aspect, world of available tags is narrowed by project. For example, only tags associated with an active project indicated by a user (or some other indication, such as location) are available for display. Contextual information associated with the tag can be also be used to select the tag. Contextual information is used by matching current contextual information with the contextual information associated with the tag,” (page 5, para [0046]-[0048]).).
Jhawar in view of Pathak is not relied upon teaching but Plantinga teaches comparing the at least one tag is based on a similarity calculation (“…Where the automated platform 102 determines that, during an initial attempt to select tags from tags 260 to apply to a particular video, none of the tags in tags 260 matches or corresponds to the title and/or summary 258-n of that video (e.g., a similarity score between each tag in tags 260 and the content in the title and/or summary 258-n is less than a first threshold), the automated platform 102 may cause the LLM 104 to re-generate a longer summary ...,” (pages 4-5, para [0035]). The similarity calculation includes computing a similarity score and / or comparing the similarity score to a threshold.).
Before the effective filling date of the claimed invention, it would have been obvious to one having ordinary skill in the art to apply the teachings of Plantinga to Jhawar in view of Pathak. The motivation would have been to quantify how close a tag is to a desired tag. Additional motivation would have been to enable a consistent quality metric and / or quality control for tags. Additional motivation would have been to ensure that final tags are accurate and relevant. Additional motivation would have been to improve the accuracy and / or relevancy of tags. Additional motivation would have been to enable tag creation when a user command is unclear / uncertain.
Claim 13 is rejected under 35 U.S.C. 103 as being unpatentable over Jhawar in view of Pathak in further view of Tomar et al. (US 20180358005 A1; hereinafter Tomar).
Regarding claim 13, Jhawar in view of Pathak teaches the computer system of claim 11, wherein detecting the at least one annotation caption comprises analyzing the audio data using at least one of a acoustic model or a language model (Jhawar; “At step 820, audio analysis is performed on the audio signal to identify a tag initiation command issued by a user of the head-mounted display device. The tag initiation command comprises a tag activation word and a tag description,” (page 7, para [0078]).
“The audio processing component 216 processes an audio signal and recognizes human speech within the audio signal. The human speech can be converted to text by the audio processing component. The text can then be used to control tagging functions,” (Jhawar; page 4, para [0042]).
“The audio signal can be processed with an acoustic model, which identifies sounds within the audio signal. The sounds can then be processed by a language model, which matches the sounds to words, phrases, and or sentences,” (Jhawar; page 4, para [0043]).).
Jhawar in view of Pathak is not relied upon teaching but Tomar teaches analyzing the audio data using at least one of a long short-term memory (LSTM) model or a large language model (LLM) (“a speech to text ASR optionally combined with a natural language processing (NLP) system to either map the input acoustic signal to one of the target outcomes intended by the user with or without a level of confidence for this mapping, or to transcribe the input acoustic signal to text in the desired language of the user, wherein such a speech recognition system can be pre-trained using any one or more of acoustic modeling techniques, such as HMMs, GMMs, DNNs, CNN, RNNs, LSTM, GRU, HAC, etc,” (page 2, para [0025]).
“The ASR system includes one or more acoustic models trained using any one or more of acoustic modeling techniques such as HMMs, GMMs, DNNs, RNNs, LSTM, GRU, HAC, etc., possibly combined with a NLP module to map the recognized text to one of the intended actions,” (page 2, para [0029]; page 1, para [0010]).).
Before the effective filling date of the claimed invention, it would have been obvious to one having ordinary skill in the art to apply the teachings of Tomar to Jhawar in view of Pathak. The motivation would have been to “provide a high degree of recognition accuracy,” (page 2, para [0022]). Additional motivation would have been to improve the recognition accuracy. Additional motivation would have been to improve computational efficiency and / or speed.
Claims 5 and 6 are rejected under 35 U.S.C. 103 as being unpatentable over Jhawar in view of Pathak in further view of Furuta et al. (US 20130216058 A1; hereinafter Furuta).
Regarding claim 5, Jhawar in view of Pathak are not relied upon teaching but Furuta teaches the computer-implemented method of claim 1, wherein determining the at least one annotation interval comprises isolating an interval containing a defect sound by performing a frequency analysis on the audio data (“the noise suppression device comprises: the Fourier transform unit 2 for converting the input signal in the time domain to the spectral components in the frequency domain; the power spectrum calculation unit 3 for calculating the power spectrum from the spectral components; the voice/noise section decision unit 4 for deciding the noise section of the input signal; the noise spectrum estimation unit 5 for estimating the noise spectrum from the input signal in the noise section,” (page 7, para [0098]; Fig 1).
“sets the decision flag Vflag at "1 (voice)" as voice when at least one of the following Expressions (3) and (4) are satisfied, and sets the decision flag Vflag at "0 (noise)" as noise in the other cases,” (page 2, para [0032]-[0033]).
“S.sub.pow and N.sub.pow are the sum total of the power spectrum of the input signal and the sum total of the estimated noise spectrum, respectively,” (pages 2-3, para [0034]; page 2, Expressions (3) and (4)).
Isolating an interval includes deciding whether the input signal of the present section / frame is voice or noise. The defect sound includes noise and / or noise section. Performing a frequency analysis on the audio data includes using the Fourier transform unit 2 for converting the input signal in the time domain to the spectral components in the frequency domain.).
Before the effective filling date of the claimed invention, it would have been obvious to one having ordinary skill in the art to apply the teachings of Furuta to Jhawar in view of Pathak. The motivation would have been “for improving the recognition rate of a voice recognition system,” (page 8, para [0111]). Additional motivation would have been to improve the accuracy of voice recognition and / or improve the accuracy of detecting user commands. Additional motivation would have been to enable voice commands in noisy environments.
Regarding claim 6, Jhawar in view of Pathak are not relied upon teaching but Furuta teaches the computer-implemented method of claim 5, wherein performing the frequency analysis comprises refining the isolated interval by performing an audio smoothing technique, the audio smoothing technique comprising at least one of moving average technique, a median filtering technique, or a Gaussian smoothing technique ((page 7, para [0098]; Fig 1). “To correct the estimated noise spectrum, a median filter as shown in the following Expression (9) is used, for example, and the filter is switched in accordance with the magnitude of the variance V(.lamda.). Incidentally, the term "median filter" refers to the processing of rearranging signals in a prescribed region in the order of power and of smoothing by taking its median,” (pages 3-4, para [0051]-[0052]).
“when the decision flag Vflag=1, since the present frame is voice, the smoothed estimated noise spectrum N (.lamda.-1,k) obtained by the previous frame is output. This makes it possible to stop excessive smoothing, and to prevent the voice signal erroneously mixed into the estimated noise spectrum from having an effect on the correction spectrum, thereby being able to carry out good noise suppression,” (page 4, para [0055]-[0056]).
“as the correction processing of the estimated noise spectrum, it is possible to execute at least one of the frequency direction smoothing and interframe smoothing,” (page 7, para [0100]).
“When the decision flag Vflag=0 in the foregoing Expression (7), since the input signal of the present frame is decided as noise, the noise spectrum estimation unit 5 updates the estimated noise spectrum N(.lamda.-1,k) of the previous frame,” (page 3, para [0040]-[0044]).
Refining the isolated interval includes correcting the estimated noise spectrum. A median filtering technique includes Furuta’s use of a median filter. The isolated interval includes the estimated noise spectrum.).
Before the effective filling date of the claimed invention, it would have been obvious to one having ordinary skill in the art to apply the teachings of Furuta to Jhawar in view of Pathak. The motivation would have been “for improving the recognition rate of a voice recognition system,” (page 8, para [0111]). Additional motivation would have been to improve the accuracy of voice recognition and / or improve the accuracy of detecting user commands. Additional motivation would have been to enable voice commands in noisy environments.
Claim 18 is rejected under 35 U.S.C. 103 as being unpatentable over Jhawar in view of Pathak in further view of Ge et al. (US 20180158463 A1; hereinafter Ge).
Regarding claim 18, Jhawar in view of Pathak is not relied upon teaching but Ge teaches the computer program product of claim 17, wherein determining the at least one annotation interval comprises: extracting an annotated audio range (“In operation 410, the preprocessor 172 extracts the voiced portions of the audio received from the IMR 122 (e.g., the portions containing speech rather than merely silence or line noise),” (page 6, para [0079]).
After combination, Ge’s voiced portions of the audio corresponds to Jhawar’s duration of video (Jhawar; page 2, para [0020]) and becomes the annotated audio range.); and
identifying an annotated audio interval by performing an audio smoothing technique in association with at least one acoustic feature, the at least one acoustic feature comprising at least one of an energy feature, a pitch feature, a spectral centroid feature, or a spectral bandwidth feature (“this is achieved by tuning the signal median smoothing parameters, such as step size and smoothing order, as well as setting the thresholds TE and TC as a weighted average of the local maxima in the distribution histograms of the short-term energy and spectral centroid respectively. FIGS. 5B and 5C illustrate examples of applying different median filter smoothing step sizes to Short Time Energy (STE) and Spectral Centroid (SC),” (page 7, para [0087]; Fig 5B; Fig 5C; pages 6-7, para [0084]).
“the preprocessor 172 only considers the speech frame to be voiced when E and C are both above their thresholds TE and TC and classifies the speech frames as voiced frames or unvoiced frames, accordingly. In one embodiment, in operation 420, unvoiced speech frames are removed, and the voiced speech frames are retained and output for further processing,” (page 7, para [0086]).
Identifying an annotated audio interval includes classifying the speech frames as voiced frames or unvoiced frames and / or retaining the voiced and removing the unvoiced speech frames. An annotated audio interval includes voiced speech frames. At least one acoustic feature includes the short-term energy and / or spectral centroid.).
Before the effective filling date of the claimed invention, it would have been obvious to one having ordinary skill in the art to apply the teachings of Ge to Jhawar in view of Pathak. The motivation would have been to improve the accuracy of voice recognition and / or improve the accuracy of detecting user commands. Additional motivation would have been to enable reducing CPU load and / or improving resource efficiency. Additional motivation would have been to enable voice commands in noisy environments.
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure.
Dantrey et al. (US.
Ding et al. (US 20080059163 A1) teaches the state of the art of moving average technique like that of claim 6.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to ERICA G THERKORN whose telephone number is (571)272-2939. The examiner can normally be reached Monday - Friday 9:00am - 5:00pm.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Devona Faulk can be reached at 571-272-7515. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/ERICA G THERKORN/Examiner, Art Unit 2618
/DEVONA E FAULK/Supervisory Patent Examiner, Art Unit 2618