Minh-Ha Duong, IC Mask Layout Designer at Maxim Integrated Products, San Francisco Bay Area.
Dong Minh Ha NH Binh Duong, Vietnam.
Minh Ha Duong, Commissioning Editor at Oxford University Press, Oxford, United Kingdom
Minh Ha Duong, Commissioning Editor at Folens Publishers, Oxford, United Kingdom
Duong Ha Minh, Facebooker, a étudié à DHNN-DHQGHN, Habite à : Hanoï
Ha Minh Duong Facebooker, a étudié à ĐH Tài chính - Marketing
Affichage des articles dont le libellé est incertitude. Afficher tous les articles
Affichage des articles dont le libellé est incertitude. Afficher tous les articles
vendredi 9 septembre 2011
mardi 23 février 2010
Statistique morbide
vendredi 25 septembre 2009
Conference on ambiguity, uncertainty and climate change at Berkeley
Contrairement à ce que le billet précédent pourrait laisser penser, je n'ai pas pris l'avion que pour sortir danser et manger des entrecôtes. A propos, dans les 747 à l'étage les places côté hublot sont les mieux car il y a plus de place pour les coudes au dessus des coffres à bagage situés entre le siège et la carlingue. Et oui, j'ai compensé mes émissions de CO2 chez Action Carbone.
Donc la conférence réunissait une quarantaine de chercheurs et une dizaine de doctorants de Berkeley. La majorité des spécialistes du sujet étaient là, du moins parmi les américains. Sinon des européens et Quiggin qui est australien comme tout le monde le sait. Il faut constater que quand c'est M Hanneman -un des pères de l'économie de l'environnement- qui invite, les gens intéressants viennent, donc c'est intéressant. Cercle vertueux.
Voir en particulier les présentations de:
Je crois qu'il y a consensus sur l'utilisation des multiple priors pour modéliser l'ambiguité (langage d'économistes) c'est à dire les probabilités imprécises pour représenter l'incertitude (en langage de statisticien). L'idée d'abandonner l'axiome de complétude des préférences (il serait temps !) a été entendue brièvement en séance et dans les couloirs, mais probablement pas par tout le monde. Ditto que la somme des probabilités peut être inférieure à 1 quand on doute trop.
Concernant le problème de la double incertitude, le pendant économique du paramètre de la sensibilité climatique est la valeur du CO2 au dessus de laquelle on peut supposer une technologie de mitigation illimitée (coût de la backstop / choke price).
Les modélisateurs intégrés et économiques sont-ils toujours miséreux par rapport aux climatophysiciens qui eux sont de vrai scientifiques et valident leurs modèles sur 200 ans de données en arrière avant de faire les projections ? La question est posée.
George Berkeley était un philosophe irlandais, il a donné son nom à un lieu charmant en corniche surplombant l'ouverture de la Baie vers le Pacifique et l'Asie (avec le Golengate Bridge) où s'est établie l'Université de Californie. On l'appelle Berkeley depuis que UC est devenu un système, mais l'équipe de foot reste les Cal Bears.
Donc la conférence réunissait une quarantaine de chercheurs et une dizaine de doctorants de Berkeley. La majorité des spécialistes du sujet étaient là, du moins parmi les américains. Sinon des européens et Quiggin qui est australien comme tout le monde le sait. Il faut constater que quand c'est M Hanneman -un des pères de l'économie de l'environnement- qui invite, les gens intéressants viennent, donc c'est intéressant. Cercle vertueux.
Voir en particulier les présentations de:
- Ani Guerdjikova de Cornell sur l'apprentissage statistique à partir d'un nombre fini d'expériences avec le modèle de Shafer.
- Sujoy Mukerji sur la version récursive du modèle KMM,
- Christian P. Traeger sur les préférences récursives assez proche de mon article avec Nicolas Treich,
- le géostatisticien Stephan Sain à UCAR, parce que j'ai besoin de géostatistique pour travailler sur les analogues climatiques,
- le psychologue expérimental David Budescu, déjà rencontré à l'EPA, Washington, DC il y a 2 ans, parce que j'ai besoin de psychologie expérimentale pour tester les modèles de prise de décision en probabilités imprécises,
- John Quiggin qui propose une interprétation du principe de précaution formalisée mais innovante. Il m'a expliqué pourquoi son blog team s'appelle Crooked Timber, c'est une citation de Kant. Son compte rendu de le conférence est plus politiquement substantiel que le mien (je crois qu'il a aussi 10 ans d'expérience de blogger en plus).
Je crois qu'il y a consensus sur l'utilisation des multiple priors pour modéliser l'ambiguité (langage d'économistes) c'est à dire les probabilités imprécises pour représenter l'incertitude (en langage de statisticien). L'idée d'abandonner l'axiome de complétude des préférences (il serait temps !) a été entendue brièvement en séance et dans les couloirs, mais probablement pas par tout le monde. Ditto que la somme des probabilités peut être inférieure à 1 quand on doute trop.
Concernant le problème de la double incertitude, le pendant économique du paramètre de la sensibilité climatique est la valeur du CO2 au dessus de laquelle on peut supposer une technologie de mitigation illimitée (coût de la backstop / choke price).
Les modélisateurs intégrés et économiques sont-ils toujours miséreux par rapport aux climatophysiciens qui eux sont de vrai scientifiques et valident leurs modèles sur 200 ans de données en arrière avant de faire les projections ? La question est posée.
George Berkeley était un philosophe irlandais, il a donné son nom à un lieu charmant en corniche surplombant l'ouverture de la Baie vers le Pacifique et l'Asie (avec le Golengate Bridge) où s'est établie l'Université de Californie. On l'appelle Berkeley depuis que UC est devenu un système, mais l'équipe de foot reste les Cal Bears.
vendredi 11 juillet 2008
Complément d'observation sur les ensembles disjonctifs
Au début de la théorie de probabilité*, on pose un ensemble Ω qui est souvent appelé ensemble des états du monde possibles. Je préfère l'appeler cadre de référence (vocabulaire de la théorie de l'évidence) ou ensemble des états du monde descriptibles (merci en passant à Pierre M. qui m'a expliqué la différence entre possiblité et dicibilité quand on parlait de modélisation). Bref, on suppose que les éléments de Ω décrivent des états du monde mutuellement exclusifs: si la pièce tombe sur pile elle ne tombe pas sur face (on suppose aussi qu'ils sont collectivement exhaustif: la pièce ne tombe pas sur la tranche.)
Dans CONSTRAINTS ON WORD LEARNING: SPECULATIONS ABOUT THEIR NATURE, ORIGINS, AND DOMAIN SPECIFICITY, Ellen Markman explique que l'hypothèse d'exclusivité mutuelle est naturelle. Elle aide les enfants à comprendre ce qu'un nouveau mot désigne, quand ils le rencontrent pour la première fois, en leur permettant de procéder par élimination.
La notion de classe vient après, en passant par celle de collection.
*Note: Sauf pour la théorie DSmT, mais qui prétend que du jaune ET du bleu c'est possible ça fait du vert...
Dans CONSTRAINTS ON WORD LEARNING: SPECULATIONS ABOUT THEIR NATURE, ORIGINS, AND DOMAIN SPECIFICITY, Ellen Markman explique que l'hypothèse d'exclusivité mutuelle est naturelle. Elle aide les enfants à comprendre ce qu'un nouveau mot désigne, quand ils le rencontrent pour la première fois, en leur permettant de procéder par élimination.
La notion de classe vient après, en passant par celle de collection.
*Note: Sauf pour la théorie DSmT, mais qui prétend que du jaune ET du bleu c'est possible ça fait du vert...
mardi 8 juillet 2008
ISIPTA final day 5
Bayesian robustness, by Fabrizio Ruggeri
Basics: To do Bayesian analysis of some data, you need (a) to select a parametric Model, (b) some Priors and (c) a Loss function to optimize over.
To test the model goodness after the analysis: visual fits, tests like Khi 2 or KS are done even if not casher. The real Baysian thing is to look at Bayes factors BF = f1(x) / f2(x), which can be used to select the model that maximizes the probability of the observed data. Posterior odds are also Baysian, but not as popular as BF, they are too sensitive to the priors.
Personally, I would rather use frequencies than improper priors.
Giron and Rios (1980) paper is famous:
If the preferences satisfy some axioms, then
There is a set of prior probabilities such that
b is preferred to a iff the expected value of the loss with a is >= ev loss with b FOR ANY probability in the set
Historical notes and references:
Bayesian robustness was booming in early 90's and busting by mid 90. The new new thing now for Bayesians is Markov Chain Monte Carlo. What remains is that people are aware of the need for multiple priors and sensitivity analysis to hyperparameters. What is missing is COTS software.
- Review paper in 1994 by Jim Berger in the Spanish Journal "Tests", with discussion.
- Proceedings of the two conferences in Spain (Valencia ?), one published in JSDI, the other in the Institute for Mathematical Statistics lecture note.
- Kadane's book.
- The handbook "Robust bayesian analysis" by Berger, Shyamalkumar and myself was perhaps the swan's song.
Glenn Shafer: What is risk? What is probability? Game-theoretic answers.
Objective vs. subjective goes on for 171 years (Simon-Denis Poisson 1837).
Our more concrete question: is there a repetitive structure for the question and the data ? By that, we mean online prediction (information feedback before you make the next prediction). Not necessarily iid or physical aleatory.
Yes => we can make good probability forecasts
No => we must weight evidence
GS 76: repetitive structure weak
GS 96: repetitive structure strong
S/Vovk 01: unifying with Game Theory
The foundations are not measure theory or Kolmogorov axioms. We start with games.
Cournot's citation "A physically impossible event is one whose probability is infinitely small. This remark alone gives substance -an objective and phenomenological value- to the mathematical theory of probability" Examples: Balancing a cone on its head. Being hit on the head by a rooftop tile. Another way to say "Small probabilities don't happen to me." is the way to connect mathematical theory with real world.
Our fundamental principle swiches that to the efficient market hypothesis: you won't multiply your capital by a large factor if you bet with the probabilities (no line of credit allowed). NB: IRL, traders get rich by betting someone else's money.
Fondations des probas par la théorie des jeux
Méthode générale:
La nature joue une séquence y = (y_i) , on n'impose aucune contrainte sur les (y_i) c'est sa stratégie.
Pour prouver une propriété P(y), i.e. la loi des grands nombres
- On construit un jeu tel que la condition de gain est:
- le capital reste toujours positif
- le joueur gagne si P ou si il devient infiniment riche - On montre qu'il existe une stratégie gagnante
- Principe fondamental: pas de martingales positives qui marchent
- Donc P
Remarques:
L'idée est que la nature peut toujours
- empêcher le joueur de s'enrichir, ou
- violer P
mais si le joueur mise astucieusement sur l'écart à P, pas les deux en même temps.
Intérêts de l'approche:
- C'est une exploration mathématique des méthodes de spéculation. Les preuves étant constructives, elles sont encore plus intéressantes si on ne croit pas au principe fondamental !
- Pas besoin de l'axiome d'additivité dénombrable ou finie.
- Preuves parfois plus courtes que dans l'approche traditionnelle, il suffit d'exhiber la stratégie.
- On introduit directement les lower et upper expectations (probabilités imprécises)
Pub: Special issue of the JEHPS e-journal.
Comment faire de bonnes prévisions
(Hot from my desk, see August 2007 Working paper #22 on defensive forecasting.)
Remark: if skeptic has a strategy A that gets infinitely rich if Pa is violated, and B that gets infinitely rich if Pb is violated, then (A+B)/2 gets infinitely rich if either Pa or Pb is violated. By averaging repeatedly, Skeptic can build a quasi-universal strategy that gets rich if nature violates any number of laws of how probability should behave. The universal test is not computable, so it is a quasi-universal test in practice.
Now take the point of view of the forecaster. The skeptic uses the fixed strategy above and plays before forecaster. We assume, critically, that the strategy of skeptic is continuous with respect to the forecaster's play at the next step.
Justifications:
- The strategies we had in the book are indeed continuous.
- Brouwer held that all constructible functions are continuous.
- Another way around is to allow forecaster imprecision.
Then there are forecasting strategies that prevent skeptic from making money.
Pub: See our Online prediction wiki.
Idea: see what happens if we allow the forecaster to make upper and lower previsions.
lundi 7 juillet 2008
ISIPTA school day 4
Oulah, les fonctions transcendantales le lundi matin ça déchire la tête. Les Gamma, Beta, BeBi, Diri, DiMn, factorielle ascendante et autres coefficients binomiaux généralisés n'ont plus de secret pour nous. Les processus de Markov non plus.
Vocabulary warning: statistical "inference" is about learning from data. Do not confuse with logical "inference" which is about reasoning.
There are 2 ways to say that the future will look like the past. One is to assume a random sampling model, the other is to assume exchangeability.
In a probability tree, nodes represent values of variables. In a bayesian or credal network, nodes represent the variables themselves.
An imprecise Markov chain can be seen as a credal net, where the graph is a chain (1 parent, 1 child), under the "epistemic irrelevance" notion of independance.
Ideas: Do experimental economics to assess empirically the hyperparameter s. Use real-world data, for example course choice by students are gambles...
What is the expected value of future information in an imprecise model ?
Make a javascript interactive page for the IDM and IDMM (people type in the counts, they get lower and upper probabilities, expectations and expected frequencies for s=1 and s=2, and n'=1 / custom. ).
Vocabulary warning: statistical "inference" is about learning from data. Do not confuse with logical "inference" which is about reasoning.
There are 2 ways to say that the future will look like the past. One is to assume a random sampling model, the other is to assume exchangeability.
In a probability tree, nodes represent values of variables. In a bayesian or credal network, nodes represent the variables themselves.
An imprecise Markov chain can be seen as a credal net, where the graph is a chain (1 parent, 1 child), under the "epistemic irrelevance" notion of independance.
Ideas: Do experimental economics to assess empirically the hyperparameter s. Use real-world data, for example course choice by students are gambles...
What is the expected value of future information in an imprecise model ?
Make a javascript interactive page for the IDM and IDMM (people type in the counts, they get lower and upper probabilities, expectations and expected frequencies for s=1 and s=2, and n'=1 / custom. ).
vendredi 4 juillet 2008
ISIPTA summer school, day 3
Aujourd'hui c'était pointu, encore plus intense en algorithmes et en définitions ésotériques. Bref on est arrivé dans la recherche en train de se faire après avoir passé les deux premiers jours à résumer Le livre, ce qui n'était pas une mince affaire car il pèse 720 pages de formules mathématiques bien serrées.
Comme tout le monde j'ai suivi de loin avec des passages en terrain connu, sur la prise de décision dans mon cas. L'après midi s'est terminé en revenant à des présentations d'applications concrêtes, c'était agréablement plus digeste. J'ai bien aimé celle de Giorgio Corani sur le Naive Credal Classifier JNCC2 qu'il a programmé avec Marco Zaffalon, l'approche est claire et marche bien. Le classificateur imprécis donne des résultats ambigus mais il est plus fiable. En d'autres termes il se trompe moins souvent, mais s'autorise plusieurs réponses.
Un problème général de ce genre de travail: il faut trouver une façon de la comparer avec l'approche précise classique qui soit vraiment juste, mais cela relève d'un arbitrage entre l'aversion aux faux positifs et l'aversion aux faux négatifs. Voilà encore une idée de papier (court) !
Après le travail nous avons eu droit à la visite guidée de Montpellier. La guide nous a enseigné quelques expressions locales: "Faire l'oeuf" (= être Place de la Comédie), "Saint Roch et son chien" (= inséparables amis, mais on ne dit pas lequel est le chien), "I can't joke in English, I am blonde !".
En sortie du vendredi soir: quelques longueurs dans la piscine olympique à côté de l'hotel, retour au centre ville pour apprécier la musique locale et les animations municipales de lancement de la saison estivale (3€ = 3 canons, et verre souvenir en plus !), et fermeture de quelques bars avec Marco.
Comme tout le monde j'ai suivi de loin avec des passages en terrain connu, sur la prise de décision dans mon cas. L'après midi s'est terminé en revenant à des présentations d'applications concrêtes, c'était agréablement plus digeste. J'ai bien aimé celle de Giorgio Corani sur le Naive Credal Classifier JNCC2 qu'il a programmé avec Marco Zaffalon, l'approche est claire et marche bien. Le classificateur imprécis donne des résultats ambigus mais il est plus fiable. En d'autres termes il se trompe moins souvent, mais s'autorise plusieurs réponses.
Un problème général de ce genre de travail: il faut trouver une façon de la comparer avec l'approche précise classique qui soit vraiment juste, mais cela relève d'un arbitrage entre l'aversion aux faux positifs et l'aversion aux faux négatifs. Voilà encore une idée de papier (court) !
Après le travail nous avons eu droit à la visite guidée de Montpellier. La guide nous a enseigné quelques expressions locales: "Faire l'oeuf" (= être Place de la Comédie), "Saint Roch et son chien" (= inséparables amis, mais on ne dit pas lequel est le chien), "I can't joke in English, I am blonde !".
En sortie du vendredi soir: quelques longueurs dans la piscine olympique à côté de l'hotel, retour au centre ville pour apprécier la musique locale et les animations municipales de lancement de la saison estivale (3€ = 3 canons, et verre souvenir en plus !), et fermeture de quelques bars avec Marco.
jeudi 3 juillet 2008
Notes on imprecise probabilities from ISIPTA summer school
The lower prevision of categorical information "x in A" is P(f) = min f(x), x in A.
The lower and upper prevision of precise probabilistic beliefs is the expectation of f.
For real variables, we also have lower and upper cumulative distribution L(x) = P( (-∞ , x] )
There are always three points of view: Lower expectations, Credal sets and Desirable gambles.
P(f|B) is defined as the willingness to pay x for a gamble that rewards f(), canceled with money back if B happens. Its definition implies a way to compute it, the Generalized Bayes Rule:
P(f|B) is the x solution of P(1B (f - x) ) ) = 0.
If P(B)>0, then P(f|B) is also the lower enveloppe of the P(f|B), for all P in the credal set, conditioned using Bayes rule : p(x,y) = p(x) p(y|x)
Full conditional measures are Probabilities coming in layers:
in layer n+1, you define p(x|y) for the y such that p(y)=0 in layer n...
and nothing blows up !
There are no probabilities, only conditional probabilities. See Krauss-Dubins representation. See Isaac Levi's (and Peter Walley's) book.
If you speak only about desirable gambles, you are limited to convex credal sets because they define half-spaces. But see this 2007 paper by Seidenfeld, Kadane and Schervish working with general partial preferences ordering, it gives meaning to non-convex credal set. In practice convexity is not so important for computation, one is always at the vertices.
The lower and upper prevision of precise probabilistic beliefs is the expectation of f.
For real variables, we also have lower and upper cumulative distribution L(x) = P( (-∞ , x] )
There are always three points of view: Lower expectations, Credal sets and Desirable gambles.
P(f|B) is defined as the willingness to pay x for a gamble that rewards f(), canceled with money back if B happens. Its definition implies a way to compute it, the Generalized Bayes Rule:
P(f|B) is the x solution of P(1B (f - x) ) ) = 0.
If P(B)>0, then P(f|B) is also the lower enveloppe of the P(f|B), for all P in the credal set, conditioned using Bayes rule : p(x,y) = p(x) p(y|x)
Full conditional measures are Probabilities coming in layers:
in layer n+1, you define p(x|y) for the y such that p(y)=0 in layer n...
and nothing blows up !
There are no probabilities, only conditional probabilities. See Krauss-Dubins representation. See Isaac Levi's (and Peter Walley's) book.
If you speak only about desirable gambles, you are limited to convex credal sets because they define half-spaces. But see this 2007 paper by Seidenfeld, Kadane and Schervish working with general partial preferences ordering, it gives meaning to non-convex credal set. In practice convexity is not so important for computation, one is always at the vertices.
mercredi 2 juillet 2008
Notes sur la conférence de D. Dubois introductive de l'école d'été ISIPTA (Montpellier, 2-8 juillet 2008)
Il faut distinguer la connaissance générique, qui porte sur les classes, de l'évidence singulière ): désolé pour la traduction :( qui porte sur les objects ou situations uniques. Tirer des conclusions à propos d'un cas demande de combiner les deux types d'information.
Soit un tableau de données, à chaque ligne on a observé si A s'est produit ou pas, et si B s'est produit ou pas. La croyance à propos de l'implication matérielle "si B alors A" i.e. "A si B" ou "A|B" devrait être fonction
- croissante de la proportion d'exemples "A ∩ B", mais aussi
- décroissante de la proportion de contre-exemple "(¬ A) ∩ B"
- indifférente aux autres cas "¬ B".
Remarque: deux propositions contraposées "B⇒A" et "¬A ⇒ ¬B" ont les mêmes contre exemple mais des exemples différents. Cela fait que leur niveau de certitude tiré du tableau ne seront pas nécessairement les mêmes.
La logique modale aléthique correspondant à la théorie des possibilités booléennes répond au doux nom de KD45.
Dans une logique modale on a un opérateur ☐ qui signifie "je sais". Considérons la proposition A: le président est vivant.
La formule "A ou ¬A" est une tautologie.
Mais pas la formule "☐ A" ou "☐ ¬ A" qui dit que je suis informé en haut lieu.
D'habitude un ensemble est une réunion d'objets: un service de six verres. En probabilité on utilise des ensembles disjonctifs: les valeurs sont mutuellement exclusives. En fin de compte, il n'y a qu'un seul verre qui est le bon.
C'est aussi la différence entre théorie des possibilités et théorie floue. Une distribution de possibilité est portée par un ensemble flou disjonctif.
Comme la théorie des possibilités n'est pas additive, les niveaux de possibilité peuvent prendre leurs valeurs sur un ensemble ordonné pas forcément numérique: treillis, chaîne... Cela fait qu'en pratique il y a deux branches de la théorie, la qualitative et la quantitative qui utilise [0,1].
Un pari échangeable c'est une proposition dans laquelle A demande à B de proposer son prix pour un bien contingent... en se réservant l'option d'acheter ce bien à B à ce prix au lieu de le lui vendre. Comme ça B est incité à donner un prix correct sinon il se fait dépouiller !
Faire des probabilités imprécises c'est mettre ensemble la logique et les probabilités. On peut toujours voir une BBA de deux façons: comme un ensemble généralisé (aléatoire), ou comme une probabilité généralisée (imprécise).
Là où en probabilités précises il y a un opérateur (règle Bayes par ex.), en probabilité imprécises il y a en plusieurs généralisations. Il ne s'agit pas de savoir quelle est la bonne, mais quelles sont les hypothèses justifiant l'utilisation de telle ou telle généralisation.
On s'aperçoit en particulier que la probabilité de "A si B" n'est pas la même chose que "la probabilité si B" de A. Dans le premier cas, on s'intéresse à un pari contingent à B alors que dans le second cas, on met d'abord à jour les croyances en apprenant que B est vrai.
In Evidence Theory, discounting a testimony assumes that the witness invented the information. This is not the same as assuming he lies deliberately.
Hacking's Frequency Principle (Hacking, 1965): If the objective probability of A is p then the subjective probability of A is p.
Dans le cas général l'adjoint de Pl se note Cr pour certainty.
Papier à écrire depuis 10 ans: La précaution comme l'incomplétude des préférences.
Soit un tableau de données, à chaque ligne on a observé si A s'est produit ou pas, et si B s'est produit ou pas. La croyance à propos de l'implication matérielle "si B alors A" i.e. "A si B" ou "A|B" devrait être fonction
- croissante de la proportion d'exemples "A ∩ B", mais aussi
- décroissante de la proportion de contre-exemple "(¬ A) ∩ B"
- indifférente aux autres cas "¬ B".
Remarque: deux propositions contraposées "B⇒A" et "¬A ⇒ ¬B" ont les mêmes contre exemple mais des exemples différents. Cela fait que leur niveau de certitude tiré du tableau ne seront pas nécessairement les mêmes.
La logique modale aléthique correspondant à la théorie des possibilités booléennes répond au doux nom de KD45.
Dans une logique modale on a un opérateur ☐ qui signifie "je sais". Considérons la proposition A: le président est vivant.
La formule "A ou ¬A" est une tautologie.
Mais pas la formule "☐ A" ou "☐ ¬ A" qui dit que je suis informé en haut lieu.
D'habitude un ensemble est une réunion d'objets: un service de six verres. En probabilité on utilise des ensembles disjonctifs: les valeurs sont mutuellement exclusives. En fin de compte, il n'y a qu'un seul verre qui est le bon.
C'est aussi la différence entre théorie des possibilités et théorie floue. Une distribution de possibilité est portée par un ensemble flou disjonctif.
Comme la théorie des possibilités n'est pas additive, les niveaux de possibilité peuvent prendre leurs valeurs sur un ensemble ordonné pas forcément numérique: treillis, chaîne... Cela fait qu'en pratique il y a deux branches de la théorie, la qualitative et la quantitative qui utilise [0,1].
Un pari échangeable c'est une proposition dans laquelle A demande à B de proposer son prix pour un bien contingent... en se réservant l'option d'acheter ce bien à B à ce prix au lieu de le lui vendre. Comme ça B est incité à donner un prix correct sinon il se fait dépouiller !
Faire des probabilités imprécises c'est mettre ensemble la logique et les probabilités. On peut toujours voir une BBA de deux façons: comme un ensemble généralisé (aléatoire), ou comme une probabilité généralisée (imprécise).
Là où en probabilités précises il y a un opérateur (règle Bayes par ex.), en probabilité imprécises il y a en plusieurs généralisations. Il ne s'agit pas de savoir quelle est la bonne, mais quelles sont les hypothèses justifiant l'utilisation de telle ou telle généralisation.
On s'aperçoit en particulier que la probabilité de "A si B" n'est pas la même chose que "la probabilité si B" de A. Dans le premier cas, on s'intéresse à un pari contingent à B alors que dans le second cas, on met d'abord à jour les croyances en apprenant que B est vrai.
Final tidbits
En théorie des possibilités c'est l'information négative qui est intéressante, le résultat donne ce qui est impossible.In Evidence Theory, discounting a testimony assumes that the witness invented the information. This is not the same as assuming he lies deliberately.
Hacking's Frequency Principle (Hacking, 1965): If the objective probability of A is p then the subjective probability of A is p.
Dans le cas général l'adjoint de Pl se note Cr pour certainty.
Papier à écrire depuis 10 ans: La précaution comme l'incomplétude des préférences.
vendredi 27 juin 2008
IPMU final Half-day 5
Serafín Moral, University of Granada,
did the keynote on desirable gambles: "During the last few years I did lots of data mining, but I have not found the gold yet, so instead I will explain to you the last three chapters of Peter Walley's seminal 1991 book that you did not bother to read or understand."Note: I am paraphrasing, but he is 100% right on the mark there - to me at last. BTW the book is available again, it seems they have done a reprint.
Sets of desirable gambles (SDG) and partial preference ordering are the simplest and the most general mathematical model of uncertainty (Walley, 2000).
Definitions
Let X be a variable taking values in Ω finite
A gamble is a reward/loss function f: Ω -> R
D is a SDG iff
A1: if f(ω) > 0 for any ω, then f is in D
A2: for any f in D, for any a > 0, the gamble a.f is in D
A3: if f1 and f2 are in D, then f1+ f2 is in D
A SDG D avoids partial loss iff 0 is not in D
A SDG D is coherent iff it is closed and it avoids partial loss
D is a set of almost desirable gambles (SADG) iff
D is a SDG, and
A4: if f+ε in D for any ε>0, then f is in D
D avoids sure loss iff (look in the book !)
There are SADG that are not SDG.
An SDG D defines a credal set K = {P, EP( f) >= 0, for any f in D}
If D' is a SADG associated with an SDG D, then they define the same K
A credal set K defines a SADG { f: EP( f) >=0, for any P in K}
Which shows that credal sets are less general than SDG.
A SDG D defines a lower prevision WTA and upper prevision WTP:
WTA = sup {a: f-a in D}
WTP = inf {a: - f+a in D}
Conditioning
There are two notions of Conditioning a SDG D with respect to the subset B of W.
One way is to consider the SDG
D_B = { f: f.1B in D} U { f: f>0} where 1B is the indicator function of B.
Note that there is no restriction on conditioning on subsets of probability zero.
The other way is to use the credal set (loosing information)
D"={ f: EP( f)>=0 for any P in K, and(?) such that there is P in K, EP( f) > 0}
The corresponding credal set is the K"={P(.|B), P in K, P(B)>0}
Implementation
A SADG can be represented as the closure of a finite set of gambles.
Generally a SDG cannot, but there is an ε-set representation:
{ f1 + ε 1B1, ... , fn + ε 1Bn}
This is enough for today, the rest was about conditional probabilities representation, consistency checking with LP (Linear Programming), inference (given D, is f desirable ?), combination and marginalization. There are two ways to define independance, epistemic and stochastic, and that seems to be a deep problem for graphical computations.
Other talks
Hwang wants your smartphone to know what you are doing (presumably to push you advertisements...). So he needs to fusion lots of data GPS, call/sms logs, device state logs, media player actions, acceleration (?), weather (taken from web)... Lucky us there is not yet enough CPU power to do so.
Fayad works for (PSA|Renault) to detect AND recognize pedestrians at carpettime minus 2s. They fusion laser, mono cam, stereo cam and radar images to track targets. I don't want to be involved in the front-end testing !.
Daniel looked at DSmT theory, in spite of the authors character. I take out two messages: 1/ The hyperpower set can be seen as a part of the power power set P(P(W)). 2/ It does not have the 'complement' operation, so singletons are not adressable.
jeudi 26 juin 2008
Day 4 at IPMU'08
Prade on Responsability judgements. There are 3 approaches to causality in the recent litterature:
There are also 3 meanings of responsibility:
We use 1. Consider time indexes t < t'
Let us assume that an agent learns of the sequence ¬Bt, At, Bt’.
Let us call C (the context) the conjunction of all other facts known by, or reported to the agent at time t’ > t.
Let us have a nonmonotonic consequence relation |≈ (and its negation |/=) that defines what is "normal"
Definitions of Facilitation:
if the agent believes that C |≈ ¬B, and that C ∧ A |/≈ ¬B, the agent will perceive A as having facilitated the occurrence of B in context C.
Definitions of Causation:
if the agent believes that C |≈ ¬B, and that C ∧ A |≈ B, the agent will perceive A as being the cause of B in context C.
Definition of Prevention:
A prevents B if A causes ¬B.
A prevented B to take place in the reported sequence ¬Bt, At, ¬Bt’ if
i) C |/≈ ¬B (i.e. ¬B does not persist by itself)
ii) C ∧ A |≈ ¬B.
In such a case, having ¬B initially was not particularly expected, i.e. was not normal, and once A took place, having ¬B was normal. Note that the condition (i) covers two situations: either C |≈ B (and ¬Bt is exceptional), or C |/≈ B (and ¬Bt is contingent). Doing A prevents to have B becoming true by the normal course of
things in the first case, and by accident in the second case (up to the potential failure of A w. r. t. its expected consequence).
Responsibility is a second level notion. Direct responsibility is causality without coercion from another agent.
Definition of indirect responsibility:
An agent a , free from the coercion of any other agent, is perceived as indirectly responsible that B took place in context C, if
- in case of a reported sequence ¬Bt, ¬Ht, Bt’ (t’ > t), a could have prevented B to take place by performing H, provided that C ∧ H |≈ ¬B;
- in case of a reported sequence ¬Bt, Ht, Bt’ (t’ > t), a could have prevented B to take place by not performing H if C ∧ ¬H |≈ ¬B, or by performing an act A that annuls the effect of H, i.e. such that C ∧ A ∧ H |≈ ¬B.
Merit/blame is a third level, it relates to deotic relations (obligations) and moral judgements (good or bad outcomes).
Prade on extracting topics in texts.
Ontologies have become available resources for identifying relations between words in a text. WordNet and EuroWordNet are widely used thesaurus-based ontologies. In WordNet there are six main relations that may hold between a pair of words (or word expressions) w and w'. They are
- w S w': w and w' are synonyms;
- w G w': w is in the glossary definition of w';
- w I w': w specializes w' ("is-a" relation); then w' I-1 w reads w' generalizes w;
- w P w': w is a part of w; conversely w' P-1 w reads w' is composed of w;
- w D w': w and w' are in the same domain;
- w R w': w is related to w'.
Relations S, D, and R are symmetrical, while G, I and P are anti-symmetrical. Besides, S, I and D are transitive.
The "bag of words" approach to information retrieval (basic weighted keywords) suffers from the so-called keyword barrier due to term ambiguity and
vocabulary mismatch: documents are not retrieved if they don't contain search terms. The words in a text most liable to give information about its contents of the text are those that are i) frequent, or ii) are in relation of some type with many words in the text, and iii) that are sufficiently specific.
Il y avait une session sur l'emotional computing. Pitt et son team à Imperial College veulent mettre un peu d'émotions dans le cyberspace: apprendre la Honte aux agents, mettre au point des capteurs d'émotion pour mieux brancher les avatars... Trivino a équipé un travailleur avec un accéléromètre (pas un iPhone ni un Freerunner hélas) et deux électrodes digitales (pour une fois qu'on peut utiliser le mot correctement) pour mesurer la conductivité de la peau. Les données vont dans un Personalized Cognitive Assistant (controleur flou) qui prévient quand il faut faire une Pause. Tout ça pour ça ! Bon ce n'est qu'un début.... Lesot (LIP6) fait du contrat pour JMM sur l'évaluation des jeux vidéos et compte aussi électrodifier les gametesters, mais c'est pas encore fait.
DevilliersReal-life emotions detection on Human-Human spoken dialogs.
A présenté un travail analysant 20 heures d'enregistrement ~700 locuteurs. Tout est annoté (deux fois par 2 codeurs différents bien sûr), 30% des productions sont colorées émotionnellement: c'est un call center hospitalier en France. Les 5 émotions principales: Peur, Colère, Tristesse, Soulagement, Neutre. En laissant de côté la sémantique (c'est utile aussi voir, mais autre papier), on cherche à reconnaitre les émotions à partir les caractéristiques paralinguistiques du signal: énergie, rythme, fréquence, respiration (il y a 80 indicateurs, mais 25 vont bien). Résultat: ~80% de détection correcte quand on se limite à 2 classes d'émotions, ~50% pour 5. At least, the hotline bots will properly get the angstlevel of the unsatisfied customers. Thanks CNRS and the HUMANE Network of Excellence.
Pensée pour les organisateurs: Idée de fonction à intégrer dans les progiciels d'organisation de conférences: une méthode de clustering (floue évidement) sur les résumés des présentations pour faire les sessions mieux centrées thématiquement.
- Schafer 96 introduced a logical language for describing temporal trees of events in terms of five basic relations between events.
- Pearl 00 proposes a modeling of causality in systems described by structural equations, taking advantage of the idea of intervention.
- We use a non-monotonic logic representation of what is the normal course of things for an agent, leading to view the potential cause(s) of a reported change as an abnormal (conjunction of) event(s) in a given context, in agreement with cognitive psychology findings.
There are also 3 meanings of responsibility:
- Something bad happened, and you caused it or could have prevented it.
- Obligation or moral duty to report or explain your actions or someone else’s action to a given authority. Ex: Parents.
- Position, which enables you to make decisions in a given organization but implies that you must be prepared to justify your actions. Ex: President
We use 1. Consider time indexes t < t'
Let us assume that an agent learns of the sequence ¬Bt, At, Bt’.
Let us call C (the context) the conjunction of all other facts known by, or reported to the agent at time t’ > t.
Let us have a nonmonotonic consequence relation |≈ (and its negation |/=) that defines what is "normal"
Definitions of Facilitation:
if the agent believes that C |≈ ¬B, and that C ∧ A |/≈ ¬B, the agent will perceive A as having facilitated the occurrence of B in context C.
Definitions of Causation:
if the agent believes that C |≈ ¬B, and that C ∧ A |≈ B, the agent will perceive A as being the cause of B in context C.
Definition of Prevention:
A prevents B if A causes ¬B.
A prevented B to take place in the reported sequence ¬Bt, At, ¬Bt’ if
i) C |/≈ ¬B (i.e. ¬B does not persist by itself)
ii) C ∧ A |≈ ¬B.
In such a case, having ¬B initially was not particularly expected, i.e. was not normal, and once A took place, having ¬B was normal. Note that the condition (i) covers two situations: either C |≈ B (and ¬Bt is exceptional), or C |/≈ B (and ¬Bt is contingent). Doing A prevents to have B becoming true by the normal course of
things in the first case, and by accident in the second case (up to the potential failure of A w. r. t. its expected consequence).
Responsibility is a second level notion. Direct responsibility is causality without coercion from another agent.
Definition of indirect responsibility:
An agent a , free from the coercion of any other agent, is perceived as indirectly responsible that B took place in context C, if
- in case of a reported sequence ¬Bt, ¬Ht, Bt’ (t’ > t), a could have prevented B to take place by performing H, provided that C ∧ H |≈ ¬B;
- in case of a reported sequence ¬Bt, Ht, Bt’ (t’ > t), a could have prevented B to take place by not performing H if C ∧ ¬H |≈ ¬B, or by performing an act A that annuls the effect of H, i.e. such that C ∧ A ∧ H |≈ ¬B.
Merit/blame is a third level, it relates to deotic relations (obligations) and moral judgements (good or bad outcomes).
Prade on extracting topics in texts.
Ontologies have become available resources for identifying relations between words in a text. WordNet and EuroWordNet are widely used thesaurus-based ontologies. In WordNet there are six main relations that may hold between a pair of words (or word expressions) w and w'. They are
- w S w': w and w' are synonyms;
- w G w': w is in the glossary definition of w';
- w I w': w specializes w' ("is-a" relation); then w' I-1 w reads w' generalizes w;
- w P w': w is a part of w; conversely w' P-1 w reads w' is composed of w;
- w D w': w and w' are in the same domain;
- w R w': w is related to w'.
Relations S, D, and R are symmetrical, while G, I and P are anti-symmetrical. Besides, S, I and D are transitive.
The "bag of words" approach to information retrieval (basic weighted keywords) suffers from the so-called keyword barrier due to term ambiguity and
vocabulary mismatch: documents are not retrieved if they don't contain search terms. The words in a text most liable to give information about its contents of the text are those that are i) frequent, or ii) are in relation of some type with many words in the text, and iii) that are sufficiently specific.
Il y avait une session sur l'emotional computing. Pitt et son team à Imperial College veulent mettre un peu d'émotions dans le cyberspace: apprendre la Honte aux agents, mettre au point des capteurs d'émotion pour mieux brancher les avatars... Trivino a équipé un travailleur avec un accéléromètre (pas un iPhone ni un Freerunner hélas) et deux électrodes digitales (pour une fois qu'on peut utiliser le mot correctement) pour mesurer la conductivité de la peau. Les données vont dans un Personalized Cognitive Assistant (controleur flou) qui prévient quand il faut faire une Pause. Tout ça pour ça ! Bon ce n'est qu'un début.... Lesot (LIP6) fait du contrat pour JMM sur l'évaluation des jeux vidéos et compte aussi électrodifier les gametesters, mais c'est pas encore fait.
DevilliersReal-life emotions detection on Human-Human spoken dialogs.
A présenté un travail analysant 20 heures d'enregistrement ~700 locuteurs. Tout est annoté (deux fois par 2 codeurs différents bien sûr), 30% des productions sont colorées émotionnellement: c'est un call center hospitalier en France. Les 5 émotions principales: Peur, Colère, Tristesse, Soulagement, Neutre. En laissant de côté la sémantique (c'est utile aussi voir, mais autre papier), on cherche à reconnaitre les émotions à partir les caractéristiques paralinguistiques du signal: énergie, rythme, fréquence, respiration (il y a 80 indicateurs, mais 25 vont bien). Résultat: ~80% de détection correcte quand on se limite à 2 classes d'émotions, ~50% pour 5. At least, the hotline bots will properly get the angstlevel of the unsatisfied customers. Thanks CNRS and the HUMANE Network of Excellence.
Pensée pour les organisateurs: Idée de fonction à intégrer dans les progiciels d'organisation de conférences: une méthode de clustering (floue évidement) sur les résumés des présentations pour faire les sessions mieux centrées thématiquement.
mercredi 25 juin 2008
Day 3 at IPMU
Enrique H. Ruspini (Artificial Intelligence Center, Stanford Research International)
Le titre de sa conférence pleinière promettait monts et merveilles, mais la présentation était surtout drivée par les réalisations techniques.
He said that the autonomous navigation robot problem is done (Roomba !), now we are working on the autonomous swarm (= multiagent robot teams =IRL= UCAV flock = helicopters). So far we can dump 100 robots in an empty army office floor and they will manage to find a pink ball without bumping into themselves.
Theoretically: It's all about similarity. Situations (i.e., solutions) are similar if they have the same importance from every relevant viewpoint. Measures of preference, cost, utility provide the bases to define similarities. Where do similarities come from ? Possibilities can be used to define similarity. If P is a fuzzy preference relation, S(w,w')=min(P(w,w'),P(w',w)) is a similarity relation.
Multicriteria decision making: When we have U1, .. , Un utility functions : W->[0,1] (W: states) we combine them with a logic function
U=(U1 and not U2) ... (U5 or U9)
Then P(w,w')=U(w) % U(w') where % is an inverse norm operator (?)
Basic instincts are the 4Fs of Feeding, Fleeing, Fighting and Reproduction.
The silver moth is well studied animal, noticeably for its very simple brain.
Hryniewicz on Verification of Kyoto protocol: a fuzzy approach
Inventory rules are critical for the allocation of tradeable emission permits in EU, among other things. (RQ: This argument is of limited relevance, given that ETS is for big point sources which can be monitored directly).
Emissions = Sum over sectors of Activity levels times Emission factors.
We propose a mixed fuzzy-random model.
Activity levels are random, there are statistics. But emission factors have both random and nonrandom uncertainty. Expert opinion is specifically OKayed by IPCC guidelines. It has been assessed the contribute to 10-20% of total uncertainty.
Pichon on TBM Unnormalized Conjonction
This is the only combination rule that makes Belief Functions a Valuation Algebra.
It's a VA because marginalisation distributes over combination.
Computing marginals is computationally efficient in VAs.
Morillas on Fuzzy graph associated with Leontief IO table
La matrice des consommation intermédiaires X définit deux matrices:
Coefficients techniques Aij = Xij/Xj
(quantité de i utilisé pour faire une unité de j)
Coefficients de distribution Bij = Xij/Xi
(proportion de la production de i vendue à la branche j)
La matrice est creuse, et les coefficients sont très inégaux.
Le Bij est utilisé pour déterminer l'importance du Aij.
On peut ainsi faire un graphe que "Qui utilise la production de qui" dans l'économie Espagnole. En haut, les services et la construction. En bas, l'extraction des ressources naturelles.
Le titre de sa conférence pleinière promettait monts et merveilles, mais la présentation était surtout drivée par les réalisations techniques.
Technical specifications (problem) -> Theory/model -> Application ( -> Tools eventually)
New specifications --^ Revision -------^ ^-------------"
He said that the autonomous navigation robot problem is done (Roomba !), now we are working on the autonomous swarm (= multiagent robot teams =IRL= UCAV flock = helicopters). So far we can dump 100 robots in an empty army office floor and they will manage to find a pink ball without bumping into themselves.
Theoretically: It's all about similarity. Situations (i.e., solutions) are similar if they have the same importance from every relevant viewpoint. Measures of preference, cost, utility provide the bases to define similarities. Where do similarities come from ? Possibilities can be used to define similarity. If P is a fuzzy preference relation, S(w,w')=min(P(w,w'),P(w',w)) is a similarity relation.
Multicriteria decision making: When we have U1, .. , Un utility functions : W->[0,1] (W: states) we combine them with a logic function
U=(U1 and not U2) ... (U5 or U9)
Then P(w,w')=U(w) % U(w') where % is an inverse norm operator (?)
Basic instincts are the 4Fs of Feeding, Fleeing, Fighting and Reproduction.
The silver moth is well studied animal, noticeably for its very simple brain.
Hryniewicz on Verification of Kyoto protocol: a fuzzy approach
Inventory rules are critical for the allocation of tradeable emission permits in EU, among other things. (RQ: This argument is of limited relevance, given that ETS is for big point sources which can be monitored directly).
Emissions = Sum over sectors of Activity levels times Emission factors.
We propose a mixed fuzzy-random model.
Activity levels are random, there are statistics. But emission factors have both random and nonrandom uncertainty. Expert opinion is specifically OKayed by IPCC guidelines. It has been assessed the contribute to 10-20% of total uncertainty.
Pichon on TBM Unnormalized Conjonction
This is the only combination rule that makes Belief Functions a Valuation Algebra.
It's a VA because marginalisation distributes over combination.
Computing marginals is computationally efficient in VAs.
Morillas on Fuzzy graph associated with Leontief IO table
La matrice des consommation intermédiaires X définit deux matrices:
Coefficients techniques Aij = Xij/Xj
(quantité de i utilisé pour faire une unité de j)
Coefficients de distribution Bij = Xij/Xi
(proportion de la production de i vendue à la branche j)
La matrice est creuse, et les coefficients sont très inégaux.
Le Bij est utilisé pour déterminer l'importance du Aij.
On peut ainsi faire un graphe que "Qui utilise la production de qui" dans l'économie Espagnole. En haut, les services et la construction. En bas, l'extraction des ressources naturelles.
Inscription à :
Articles (Atom)