IP Library Granted Patent US 10,990,890
Granted Patent B2
US 10,990,890 · App. 16/824,025 · Granted Apr 27, 2021

Machine learning system

Inventors: Stefanos Eleftheriadis (Cambridge, GB); James Hensman (Cambridge, GB); Sebastian John (Cambridge, GB); Hugh Salimbeni (Cambridge, GB)
Assignee: SECONDMIND LIMITED
G06N7/005G06F17/17G06N3/08G06N20/00G06Q10/04G06Q50/30G06Q10/10
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,990,890
App. No.
16/824,025
Granted
Apr 27, 2021
Kind
B2
Abstract

A reinforcement learning system comprises an environment (having multiple possible states), and agent, and a policy learner.. The agent is arranged to receive state information indicative of a current environment state and generate an action signal dependent on the state information and a policy associated with the agent, where the action signal is operable to cause an environment-state change. The agent is further arranged to generate experience data dependent on the state information and information conveyed by the action signal. The policy learner is configured to process the experience data in order to update the policy associated with the agent. The reinforcement learning system further comprises a probabilistic model arranged to generate, dependent on the current state of the environment, probabilistic data relating to future states of the environment, and the agent is further arranged to generate the action signal in dependence on the probabilistic data.

Claims (25)

1. A reinforcement learning system comprising:

an environment having multiple possible states;

an agent arranged to receive state information indicative of a current state of the environment and to generate an action signal dependent on the state information and a policy associated with the agent, the action signal being operable to cause a change in a state of the environment, the agent being further arranged to generate experience data dependent on the state information and information conveyed by the action signal;

a policy learner configured to process the experience data, whereby to update the policy associated with the agent;

a probabilistic model arranged to generate probabilistic data relating to future states of the environment; and

a model learner configured to process model input data to generate the probabilistic model,

wherein:

the environment comprises a domain having a temporal dimension;

the model input data comprises data indicative of events occurring in past states of the environment;

the agent is arranged to generate the action signal further in dependence on the probabilistic data;

the probabilistic model comprises a distribution of a stochastic intensity function, wherein an integral of the stochastic intensity function over a sub-region of the domain corresponds to a rate parameter of a Poisson distribution for a predicted number of events occurring in the sub-region; and

the model learner is configured to process the model input data to generate the probabilistic model by applying a Bayesian inference scheme to the model input data, the Bayesian inference scheme comprising:

generating a variational Gaussian process corresponding to a distribution of a latent function, the variational Gaussian process being dependent on a prior Gaussian process and a plurality of randomly-distributed inducing variables, the inducing variables having a variational distribution and expressible in terms of a plurality of Fourier components;

determining, using the data indicative of events occurring in past states of the environment, a set of parameters for the variational distribution, wherein determining the set of parameters comprises iteratively updating a set of intermediate parameters to determine an optimal value of an objective function, the objective function being dependent on the inducing variables and expressible in terms of the plurality of Fourier components; and

determining, from the variational Gaussian process and the determined set of parameters, the distribution of the stochastic intensity function, wherein the distribution of the stochastic intensity function corresponds to a distribution of a quadratic function of the latent function.

2. The system of claim 1 , wherein the model learner is further configured to process the experience data generated by the agent to update the probabilistic model.

3. The system of claim 1 , wherein each of the events occurring in a past state of the environment corresponds to an occurrence of an event in a physical system.

4. The system of claim 3 , wherein the agent is associated with an entity in the physical system.

5. The system of claim 4 , wherein the reinforcement learning system is operable to send a control signal, corresponding to the generated action signal, to the entity.

6. The system of claim 3 , wherein:

the physical system comprises one or more sensors; and

the agent is arranged to receive the state information from the one or more sensors.

7. The system of claim 4 , wherein:

each of the events occurring in a past state of the environment corresponds to a taxi request in a city; and

the agent is associated with a taxi in the city.

Assignments (2)
CHANGE OF NAME Recorded Nov 6, 2020
From: PROWLER.IO LIMITED
To: SECONDMIND LIMITED
Reel/Frame 054302/0221 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 20, 2020
From: ELEFTHERIADIS, STEFANOS; HENSMAN, JAMES; JOHN, SEBASTIAN; SALIMBENI, HUGH
To: PROWLER.IO LIMITED
Reel/Frame 052175/0850 →
Priority Claims (2)
EP 17275185 · Nov 21, 2017 · regional
EP 18165197 · Mar 29, 2018 · regional
Continuity (2)
Continuation PCTEP2018077062 · Apr 10, 2018
Related Publication 20200218999A1 · Jul 9, 2020