US2025155961A1PendingUtilityA1

On demand contextual support agent with spatial awareness

Assignee: MICROSOFT TECHNOLOGY LICENSING LLCPriority: Nov 10, 2023Filed: Nov 10, 2023Published: May 15, 2025
Est. expiryNov 10, 2043(~17.3 yrs left)· nominal 20-yr term from priority
G06T 19/20G06T 17/00G06F 3/017G06F 30/10G06F 40/40G02B 2027/0174G06F 16/909G06F 3/011G02B 27/0103
57
PatentIndex Score
0
Cited by
0
References
0
Claims

Abstract

Techniques for toggling a visibility of tagged spatial data and for correlating the tagged spatial data with output of an LLM are disclosed. Scene data describing a real-world scene in which a 3D object is located is accessed. A digital file that models the 3D object is accessed. The digital file includes tagged spatial data associated with the 3D object. User input is received. The user input is directed to the 3D object. The digital file and the user input are provided as input to the LLM. The LLM correlates the user input with the tagged spatial data and generates a response. While the LLM's response is being provided to the user, a display of a hologram is toggled, where this hologram overlays at least a portion of the 3D object.

Claims

exact text as granted — not AI-modified
What is claimed is: 
     
         1 . A method for toggling a visibility of holograms and for correlating tagged spatial data with output of a large language model (LLM), said method comprising:
 accessing scene data describing a real-world scene in which a three-dimensional (3D) object is located;   accessing a digital file that models the 3D object, wherein the digital file includes tagged spatial data associated with the 3D object;   receiving, in real time and from a user, user input directed to the 3D object located in the real-world scene;   providing the digital file and the user input as inputs to an LLM, wherein the LLM is tasked with correlating the user input with the tagged spatial data of the digital file, and wherein the LLM is further tasked with generating a response to the user input; and   while the LLM's response is being provided to the user, toggling display of a hologram that overlays at least a portion of the 3D object.   
     
     
         2 . The method of  claim 1 , wherein the inputs to the LLM further include a work instruction document, and wherein the LLM's response is further based on the work instruction document. 
     
     
         3 . The method of  claim 1 , wherein the digital file includes a computer-aided design (CAD) file. 
     
     
         4 . The method of  claim 1 , wherein the user input is one or a combination of a verbal input, a gesture input, or an eye gazing input. 
     
     
         5 . The method of  claim 1 , wherein the LLM's response includes a natural language text output, and wherein the natural language text output is converted to an audio output and is played over a speaker to the user. 
     
     
         6 . The method of  claim 1 , wherein the hologram overlays only a specific part of the 3D object as opposed to overlaying an entirety of the 3D object. 
     
     
         7 . The method of  claim 1 , wherein the LLM's response is one or a combination of a visual response, an auditory response, or a haptic response. 
     
     
         8 . The method of  claim 1 , wherein the digital file of the 3D object includes a 3D asset hierarchy associated with the 3D object. 
     
     
         9 . The method of  claim 1 , wherein the LLM identifies multiple instances of the 3D object, and wherein the LLM disambiguates those multiple instances based on additional context obtained for the 3D object. 
     
     
         10 . The method of  claim 1 , wherein multiple holograms, including said hologram, are automatically generated for the 3D object based on the digital file. 
     
     
         11 . A method for toggling a visibility of holograms and for correlating tagged spatial data with output of a large language model (LLM), said method comprising:
 accessing scene data describing a real-world scene in which a three-dimensional (3D) object is located;   accessing a digital file that models the 3D object, wherein the digital file includes tagged spatial data associated with the 3D object;   receiving, in real time and from a user, user input directed to the 3D object located in the real-world scene, wherein the user input includes a gesture input, wherein a hand-pointing vector is generated based on the gesture input, and wherein the hand-pointing vector is used to determine that the user input is directed to the 3D object;   providing the digital file and the user input as inputs to an LLM, wherein the LLM is tasked with correlating the user input with the tagged spatial data of the digital file, and wherein the LLM is further tasked with generating a response to the user input; and   while the LLM's response is being provided to the user, toggling display of a visual cue that overlays at least a portion of the 3D object.   
     
     
         12 . The method of  claim 11 , wherein, as a result of providing the LLM's response and as a result of toggling the display of the visual cue, a comprehensive response is provided to the user who provided the user input, and wherein the comprehensive response includes a language aspect and a physical space aspect. 
     
     
         13 . The method of  claim 11 , wherein the LLM processes the user input to identify specific tagged spatial data from the digital file, the specific tagged spatial data being data that is determined by the LLM to correspond to one or more key terms included in the user input. 
     
     
         14 . The method of  claim 11 , wherein the LLM further consults additional curated files that are identified as being associated with the 3D object, and wherein at least one additional curated file includes a particular curated file that includes sequential usage steps for the 3D object. 
     
     
         15 . The method of  claim 11 , wherein the LLM further consults additional curated files associated with the 3D object, at least one additional curated file includes a parts directory for the 3D object. 
     
     
         16 . The method of  claim 11 , wherein:
 the scene data includes data describing a second 3D object included in the real-world scene,   the scene data is also included in the inputs to the LLM, and   the LLM's response includes details on how the second 3D object is to be used in association with said 3D object, said usage being a part of a procedural work process.   
     
     
         17 . The method of  claim 11 , wherein curating information for the real-world scene is performed based on at least one of the scene data, the digital file, or manual input. 
     
     
         18 . A computer system for toggling a visibility of tagged spatial data and for correlating the tagged spatial data with output of a large language model (LLM), said computer system comprising:
 a processor system; and   a storage system comprising instructions that are executable by the processor system to cause the computer system to:
 access scene data describing a real-world scene in which a three-dimensional (3D) object is located, wherein the scene data is tagged with spatial data describing at least the 3D object; 
 receive, in real time and from a user, user input directed to the 3D object located in the real-world scene, wherein the user input includes a query related to the 3D object; 
 provide the scene data and the user input as inputs to an LLM, wherein the LLM is tasked with correlating the user input with the tagged spatial data, and wherein the LLM is further tasked with generating a response to the user input; and 
 while the LLM's response is being provided to the user, toggle display of a visual cue that overlays at least a portion of the 3D object. 
   
     
     
         19 . The computer system of  claim 18 , wherein the spatial data is generated as a result of performing object recognition. 
     
     
         20 . The computer system of  claim 18 , wherein the spatial data is manually entered by a user.

Join the waitlist — get patent alerts

Track US2025155961A1 — get alerts on status changes and closely related new filings.

We store only your email — no account needed. See our privacy policy.