Object tracking and identification using intelligent camera orchestration
Abstract
In one embodiment, an apparatus comprises a communication interface and a processor. The communication interface is to communicate with a plurality of cameras. The processor is to obtain metadata associated with an initial state of an object, wherein the object is captured by a first camera in a first video stream at a first point in time, and wherein the metadata is obtained based on the first video stream. The processor is further to predict, based on the metadata, a future state of the object at a second point in time, and identify a second camera for capturing the object at the second point in time. The processor is further to configure the second camera to capture the object in a second video stream at the second point in time, wherein the second camera is configured to capture the object based on the future state of the object.
Claims
exact text as granted — not AI-modified1 - 20 . (canceled)
21 . A system, comprising:
interface circuitry to communicate with a plurality of cameras deployed in an environment; and processing circuitry to:
coordinate, via the interface circuitry, the cameras to track an object in the environment and capture a snapshot for identifying the object based on predicted behavior of the object and positions of the cameras, wherein the predicted behavior is predicted based on actual behavior of the object captured by one or more of the cameras; and
cause identification of the object based on the snapshot captured by one or more of the cameras.
22 . The system of claim 21 , wherein the processing circuitry to coordinate, via the interface circuitry, the cameras to track the object in the environment and capture the snapshot for identifying the object based on the predicted behavior of the object and the positions of the cameras is further to:
adjust one or more settings of one or more of the cameras to capture the snapshot for identifying the object, wherein the one or more settings comprise at least one of:
a zoom setting; or
a pan setting.
23 . The system of claim 21 , wherein the processing circuitry is further to determine the predicted behavior of the object using a machine learning model, wherein the actual behavior of the object is supplied as input to the machine learning model, and wherein the machine learning model is trained to predict future object behavior based on actual object behavior.
24 . The system of claim 21 , wherein the processing circuitry to coordinate, via the interface circuitry, the cameras to track the object in the environment and capture the snapshot for identifying the object based on the predicted behavior of the object and the positions of the cameras is further to:
detect the object in a first video stream at a first time, wherein the first video stream is captured by a first camera of the plurality of cameras; determine, based on the first video stream, the actual behavior of the object; determine, based on the actual behavior of the object, the predicted behavior of the object; and configure, via the interface circuitry, a second camera to capture the object in a second video stream at a second time, wherein the second time is after the first time, and wherein the second camera is selected from the plurality of cameras based on the predicted behavior of the object and a position of the second camera.
25 . The system of claim 24 , wherein:
the object is a person; the snapshot comprises a facial snapshot captured by the second camera, wherein the facial snapshot includes a face of the person; and the processing circuitry to cause identification of the object based on the snapshot captured by one or more of the cameras is further to:
cause an identity of the person to be determined based on the facial snapshot captured by the second camera.
26 . The system of claim 24 , wherein:
the object is a vehicle; the snapshot comprises a license plate snapshot captured by the second camera, wherein the license plate snapshot includes a license plate of the vehicle; and the processing circuitry to cause identification of the object based on the snapshot captured by one or more of the cameras is further to:
cause a license plate number of the vehicle to be determined based on the license plate snapshot captured by the second camera.
27 . The system of claim 21 , wherein the actual behavior of the object indicates a current position and a direction of travel.
28 . The system of claim 21 , wherein the plurality of cameras comprise one or more smart cameras, wherein the one or more smart cameras comprise at least some of the interface circuitry and the processing circuitry.
29 . At least one non-transitory machine-accessible storage medium having instructions stored thereon, wherein the instructions, when executed on processing circuitry, cause the processing circuitry to:
coordinate, via interface circuitry, a plurality of cameras to track an object in an environment and capture a snapshot for identifying the object based on predicted behavior of the object and positions of the cameras, wherein the predicted behavior is predicted based on actual behavior of the object captured by one or more of the cameras; and identify the object based on the snapshot captured by one or more of the cameras.
30 . The storage medium of claim 29 , wherein the instructions that cause the processing circuitry to coordinate, via the interface circuitry, the cameras to track the object in the environment and capture the snapshot for identifying the object based on the predicted behavior of the object and the positions of the cameras further cause the processing circuitry to:
adjust one or more settings of one or more of the cameras to capture the snapshot for identifying the object, wherein the one or more settings comprise at least one of:
a zoom setting; or
a pan setting.
31 . The storage medium of claim 29 , wherein the instructions further cause the processing circuitry to determine the predicted behavior of the object using a machine learning model, wherein the actual behavior of the object is supplied as input to the machine learning model, and wherein the machine learning model is trained to predict future object behavior based on actual object behavior.
32 . The storage medium of claim 29 , wherein the instructions that cause the processing circuitry to coordinate, via the interface circuitry, the cameras to track the object in the environment and capture the snapshot for identifying the object based on the predicted behavior of the object and the positions of the cameras further cause the processing circuitry to:
detect the object in a first video stream at a first time, wherein the first video stream is captured by a first camera of the plurality of cameras; determine, based on the first video stream, the actual behavior of the object; determine, based on the actual behavior of the object, the predicted behavior of the object; and configure, via the interface circuitry, a second camera to capture the object in a second video stream at a second time, wherein the second time is after the first time, and wherein the second camera is selected from the plurality of cameras based on the predicted behavior of the object and a position of the second camera.
33 . The storage medium of claim 32 , wherein:
the predicted behavior of the object indicates a predicted future position of the object at the second time; and the instructions that cause the processing circuitry to configure, via the interface circuitry, the second camera to capture the object in the second video stream at the second time further cause the processing circuitry to:
determine, based on the position of the second camera relative to the predicted future position of the object, that the object is expected to be in view of the second camera at the second time; and
select the second camera to capture the object at the second time, wherein the second camera is selected from the plurality of cameras based on determining that the object is expected to be in view of the second camera at the second time.
34 . The storage medium of claim 32 , wherein:
the object is a person; the snapshot comprises a facial snapshot captured by the second camera, wherein the facial snapshot includes a face of the person; and the instructions that cause the processing circuitry to identify the object based on the snapshot captured by one or more of the cameras further cause the processing circuitry to:
determine an identity of the person based on the facial snapshot captured by the second camera.
35 . The storage medium of claim 32 , wherein:
the object is a vehicle; the snapshot comprises a license plate snapshot captured by the second camera, wherein the license plate snapshot includes a license plate of the vehicle; and the instructions that cause the processing circuitry to identify the object based on the snapshot captured by one or more of the cameras further cause the processing circuitry to:
determine a license plate number of the vehicle based on the license plate snapshot captured by the second camera.
36 . The storage medium of claim 29 , wherein the actual behavior of the object indicates a current position and a direction of travel.
37 . The storage medium of claim 36 , wherein the actual behavior of the object further indicates an orientation or a speed.
38 . A method, comprising:
coordinating, via interface circuitry, a plurality of cameras to track an object in an environment and capture a snapshot for identifying the object based on predicted behavior of the object and positions of the cameras, wherein the predicted behavior is predicted based on actual behavior of the object captured by one or more of the cameras, wherein the actual behavior indicates a current position and a direction of travel; and causing identification of the object based on the snapshot captured by one or more of the cameras.
39 . The method of claim 38 , wherein coordinating, via the interface circuitry, the cameras to track the object in the environment and capture the snapshot for identifying the object based on the predicted behavior of the object and the positions of the cameras comprises:
adjusting one or more settings of one or more of the cameras to capture the snapshot for identifying the object, wherein the one or more settings comprise at least one of:
a zoom setting; or
a pan setting.
40 . The method of claim 38 , wherein coordinating, via the interface circuitry, the cameras to track the object in the environment and capture the snapshot for identifying the object based on the predicted behavior of the object and the positions of the cameras comprises:
detecting the object in a first video stream at a first time, wherein the first video stream is captured by a first camera of the plurality of cameras; determining, based on the first video stream, the actual behavior of the object; determining, based on the actual behavior of the object, the predicted behavior of the object; and configuring, via the interface circuitry, a second camera to capture the object in a second video stream at a second time, wherein the second time is after the first time, and wherein the second camera is selected from the plurality of cameras based on the predicted behavior of the object and a position of the second camera.Join the waitlist — get patent alerts
Track US2024037951A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.